The decision context
A model can support a language without handling the terminology, tone, document conventions, cultural context, or policy boundaries of a real locale. Translation quality is also different from doing a business task natively in that language.
Multilingual evaluation should preserve results by language and task family. An overall average can hide that a candidate works well in English but fails escalation, extraction, or safety requirements elsewhere.
Why it matters
- Critical failures may cluster in a lower-volume language and disappear in a global average.
- Literal translation can change legal, financial, medical, support, or brand meaning.
- Tokenization, retrieval coverage, locale formats, and reviewer availability can alter both quality and cost.
Questions to answer before choosing
- Which languages, locales, scripts, dialects, and code-switching patterns occur?
- Is the task translation, native generation, retrieval, classification, or structured extraction?
- Who is qualified to design expected answers and review nuanced failures?
- Which errors must be reported per language rather than averaged?
A reviewable decision path
Each step should leave a record that another reviewer can inspect.
- 01
Build native cases
Use locally realistic tasks instead of translating only an English benchmark.
- 02
Lock locale rules
Define tone, date, number, currency, address, terminology, and policy conventions.
- 03
Use qualified review
Combine native-language expertise with the domain expertise needed by the task.
- 04
Report separately
Publish pass rate, critical failures, format validity, and review effort by language and task.
- 05
Retest the pipeline
Recheck model, retrieval, prompt, translation memory, and moderation changes together.
Continue through the evidence graph
These links connect the topic to at least three concrete models, tools, workflows, comparisons, or protocols.
Sources checked
Open the original pages before relying on a time-sensitive product decision.