AI & Terminology Quality - Cheat Sheet

Cheat sheet

AI & Terminology Quality

Field recipes, prompts & principles

AI is a fast junior assistant with no concept of accountability. Use it as one.

Four roles for AI - and one rule

1
ExtractorScans, proposes candidates, clusters variants, surfaces contexts.
2
CheckerCompares texts against approved resources and flags deviations for review.
3
DrafterProduces first drafts of definitions, notes, and metadata from evidence you supply.
4
Assistant to governance - never the governorThe human and the institutional framework keep the signature. Let the model hold final authority and quality becomes vulnerable to silent, scalable error.
Candidate generation ≠ candidate validation. The model proposes; the human, the termbase, and the standards decide.

The four recipes

Recipe 1 Role: Extractor

AI-assisted term extraction

Turn a corpus or document set into a structured candidate list you can review fast - without trusting a single candidate yet.

You are assisting a terminologist with TERM EXTRACTION. You are NOT deciding what is correct; you are proposing candidates for human review.

CONTEXT
- Domain: <e.g. medical devices / EU asylum law / automotive>
- Working language(s): <e.g. EN source, EL + FR targets>
- Register / audience: <e.g. expert-to-expert; regulatory>
- Source material: <paste text, or describe the attached corpus>

TASK
1. Propose up to <N> candidate terms: single-word AND multiword units, plus domain-heavy collocations and named entities.
2. For EACH candidate give, in a table:
   - candidate term (exactly as it appears)
   - part of speech / term type
   - a VERBATIM context sentence quoted from the source (in quotation marks)
   - likely variants or spelling/casing alternatives seen in the source
   - a tentative one-line gloss
   - EVIDENCE FLAG: mark [ATTESTED] if it appears verbatim in the source, or [INFERRED] if it is your own suggestion not literally present.
3. Cluster obvious variants of the same concept together.
4. Do NOT invent terms to reach the number. If you find fewer, return fewer.
5. If a candidate is ambiguous across domains, say so in a "caution" column.

OUTPUT
A single table, sorted by frequency in the source (most frequent first). No preamble, no conclusion.

Then the human: Human after the model: review candidates against the concept system and authoritative sources; merge duplicates; reject noise; promote survivors to termbase entries; record who decided and why.

Watch for: Confident page-headers and boilerplate masquerading as terms; [INFERRED] rows dressed up to look attested; over-generation when you set N too high.

Recipe 2 Role: Checker

AI-assisted terminology QA

Compare a text against your approved terminology and surface deviations - without turning QA into a witch-hunt where every variation is treated as a crime.

You are performing TERMINOLOGY QUALITY ASSURANCE. You compare a text against an APPROVED terminology resource and report deviations. You do not rewrite the text and you do not decide final outcomes.

INPUTS
- Approved terminology (preferred terms, deprecated/forbidden terms, any usage notes): <paste glossary / termbase export>
- Text under review: <paste text>
- Style constraints, if any: <casing, regional variant, formality>

TASK
For every place the text departs from the approved terminology, produce a row with:
   - location (quote the surrounding phrase)
   - the form used in the text
   - the approved/preferred form
   - deviation type: [wrong term] [deprecated form] [inconsistent rendering] [casing] [near-synonym drift] [register shift]
   - JUSTIFIED? - your assessment of whether the deviation might be contextually defensible, with a one-line reason. Be honest: some deviations are fine.
   - confidence: high / medium / low

RULES
- Do NOT flag stylistic variation that the approved resource does not actually govern.
- Flag the SAME concept rendered two different ways even if neither is in the glossary (internal inconsistency).
- If the approved resource is silent on something, say "not governed" rather than inventing a rule.
- Separate genuine errors from "the termbase may need to evolve here."

OUTPUT
A deviations table only. End with a 3-line summary: count of likely errors, count of defensible deviations, count of "termbase gap" cases.

Then the human: Human after the model: decide which deviations are errors, which are acceptable in context, and which signal that the termbase itself should be updated. The reviewer owns every verdict.

Watch for: A model that flags difference as if difference were guilt; missed internal inconsistencies because neither form was in the glossary; over-confident 'high' ratings on judgement calls.

Recipe 3 Role: Drafter

AI-assisted definitional drafting

Get a usable first-draft definition and usage note - grounded in sources you supply, with neighbouring concepts fenced off so the model cannot smooth them into one blur.

You are drafting a DEFINITION for a terminology entry. You work ONLY from the sources I provide. You must not introduce facts, standards, or references that are not in those sources.

INPUTS
- Term: <term>
- Domain & concept system: <where it sits; broader/narrower concepts>
- Approved sources (paste, do not summarise from memory): <standards excerpts, internal docs, authoritative usage examples>
- Neighbouring concepts it must be distinguished FROM: <e.g. decreto-legge vs decreto legislativo>

TASK
1. Draft a SHORT definition (one sentence, genus + differentia where possible).
2. Draft a LONGER explanatory note (2–4 sentences).
3. Provide 1–2 NON-EXAMPLES that show the boundary with the neighbouring concept(s) named above.
4. List the exact source snippet(s) each part of the definition rests on. If something rests on no provided source, mark it [UNSUPPORTED - do not use].
5. If the sources are insufficient to define the term, say so explicitly instead of guessing.

CONSTRAINTS
- No circular definitions ("a decree is a decree that...").
- Do not cite any standard, law, or publication not present in the inputs.
- Preserve distinctions; do NOT merge the term with its neighbours.

OUTPUT
Four labelled sections: Definition / Note / Non-examples / Source mapping.

Then the human: Human after the model: test the draft against conceptual boundaries and real institutional usage; strip anything marked [UNSUPPORTED]; verify every cited snippet actually says what the draft claims.

Watch for: Elegant, balanced sentences that say almost nothing; collapsed distinctions between neighbouring concepts; invented or subtly misquoted 'sources.'

Recipe 4 Role: Checker / Drafter

AI-assisted multilingual alignment

Surface mismatches in how a concept is rendered across languages and document types - while remembering that regional standards may legitimately differ.

You are helping with MULTILINGUAL TERMINOLOGY ALIGNMENT. You surface candidate mismatches across languages and document types. You do NOT decide equivalence; a human does.

INPUTS
- Concept / source term: <term + short definition>
- Renderings to compare (by language, region, document type): <paste, e.g. EN docs, FR-FR marketing, FR-CA UI, ZH-CN, ZH-TW>
- Known regional/standard constraints: <e.g. Quebec prefers "courriel"; ZH-CN 软件 vs ZH-TW 軟體>

TASK
For each language/variant, report:
   - the rendering(s) found
   - whether they are mutually consistent within that language
   - candidate cross-language mismatches, classified as:
     [likely error] [acceptable regional variant] [register/document-type difference] [possible genuine concept-system divergence]
   - a one-line reason for each classification
   - confidence level

RULES
- Do NOT treat a legitimate regional standard as an error.
- Flag where the SAME target language uses different forms in different document types.
- Where concept systems may genuinely differ between languages (not just wording), say so - do not force a 1:1 equivalence.
- Multilingual terminology is not a vending machine: never assert exact equivalence you cannot support.

OUTPUT
One table per language plus a final "cross-language" table of mismatches with classifications.

Then the human: Human after the model: decide, per case, whether it is an error to fix, a variant to record (with its region/scope), or a divergence to document at the concept level.

Watch for: Regional standards flagged as 'errors'; forced 1:1 equivalences; missed within-language inconsistency across document types.

Choosing a model: six criteria

1Language coverage

Credibly strong in YOUR working languages and registers. Test with your own texts, not the launch announcement.

2Grounding options

Can it use your termbase, glossaries, approved docs - retrieve at generation time rather than improvise? The single most important criterion.

3Controllability

Can you make it cautious, flag uncertainty, and separate “attested” from “suggested”? A model that can’t say “I’m not sure” is auditioning for a role it must not get.

4Workflow transparency

Even if the model is a black box, can YOU document prompt, sources, constraints, and review steps? If not, you have vibes.

5Data governance

For client terms, unreleased names, regulated content: where the data goes beats how pretty the interface is. Read the data-processing terms.

6Fitness for QA

For checking, you want disciplined, not theatrical. For terminology work, “creative” is not always a compliment.

Don’t choose by reading vendor pages. Run a small bake-off: take a text you know cold, build a tiny answer key (≈20 known terms, 5 planted inconsistencies), give the same task and resources to 2–3 systems, and score them against your key. Half a day beats any leaderboard - because you don’t have average problems, you have yours.

Seven principles

  1. 1Centralize terminology resources - AI cannot ground itself in chaos.
  2. 2Define validation ownership - who approves, who rejects, who updates.
  3. 3Separate candidate generation from final approval - non-negotiable.
  4. 4Embed terminology early in workflows, not as emergency repair at the end.
  5. 5Ask AI for evidence, not just answers - snippets, uncertainty flags, attested vs inferred.
  6. 6Evaluate quality with the right framework - the framework defines what counts as an error.
  7. 7Keep human judgment wherever genuine judgment is required.

References

  • ISO 704:2022 · ISO 1087:2019 · ISO 30042:2019 (TBX) · ISO 12616-1:2021
  • Cabré (1999), Terminology; Warburton (2021), The Corporate Terminologist
  • ECQA Certified Terminology Manager - Basic / Advanced / engineering (termnet.org, ecqa.org)
  • Kalai, Nachum, Vempala & Zhang (2025), Why Language Models Hallucinate, arXiv:2509.04664
  • IATE (iate.europa.eu) · EU DGT translation quality guidelines