Writing a translation review brief for native speakers

Why Open-Ended Review Fails

The default instruction to a reviewer is some version of "can you check this translation?"

What comes back is a list of things the reviewer would have phrased differently, mixed with a few genuine errors, with no indication of which is which. The recipient cannot tell whether an item is a blocking problem or a stylistic preference, so they either apply everything — which is expensive and sometimes wrong — or apply nothing.

The reviewer has done exactly what was asked. The brief was the problem.

A good brief converts review from an opinion exercise into a diagnostic one. It asks specific questions, defines what counts as an error, and requests severity, which is what makes the output actionable.

What to Send

Reviewers need context that the translated text alone does not provide.

The video itself, not just the script. Delivery, timing, pronunciation, and how text sits against visuals are all reviewable only in the finished asset. Script-only review misses an entire category of problems.

The source, so the reviewer can check against what was actually said rather than evaluating the translation in isolation.

The glossary, so reviewers are not re-deciding settled terminology questions. This alone eliminates a large share of unhelpful feedback.

The register decision — formality level, address form, tone target — with the reasoning. Reviewers who do not know a decision was made will flag it as an error.

The audience description. Who this content is for, what they know, and what they are expected to do with it.

The content type and purpose. Instructional content, marketing content, and formal training have different standards, and a reviewer applying the wrong one produces misdirected feedback.

Known constraints. Timing limits, character limits, or anything that explains why the translation is phrased as it is. A reviewer who suggests a longer phrasing without knowing it will not fit has wasted their effort and yours.

The Questions to Ask

Replace "check this" with a specific list. A workable set for general content:

Is any of this factually wrong relative to the source? Meaning changed, content omitted, content added, or anything reversed. This is the highest-severity category and should be asked first.

Does the terminology match the glossary, and is it what a practitioner in this field would actually say? Two distinct questions in one — glossary compliance and real-world usage — and both matter.

Is the register right and consistent? Formality, tone, and address form, held throughout rather than drifting.

Does anything read as translated rather than written? Word order that follows the source language, idioms rendered literally, or constructions that are grammatically correct and unnatural. This is the question that most reliably surfaces the difference between adequate and good.

Would anything not land culturally? Examples, references, humour, or framing that would confuse or misfire with this audience.

Are numbers, dates, units, and names correct for this market? Format conventions and factual accuracy of any figure.

For audio content, add:

Is anything mispronounced? Proper nouns, technical terms, and borrowed words specifically.

Does the delivery sound natural? Emphasis, intonation, pacing, and anything that sounds synthetic in a way that distracts.

Does the audio match the video? Timing, and particularly whether it holds at the end.

For subtitles, add:

Is anything hard to read at the speed it appears?

Do the line breaks fall in sensible places?

Does the text render correctly? Characters, diacritics, and any script-specific concerns.

Severity Ratings

Ask for a severity on every item, using a small scale with clear definitions.

Blocking: changes meaning, is factually wrong, creates legal or safety risk, or makes the content unusable. Must be fixed before publication.

Quality: makes the content read as foreign, unprofessional, or careless. Should be fixed, and the release decision depends on how many there are.

Preference: the reviewer would have phrased it differently, and the existing version is not wrong. Record but do not necessarily act.

Reviewers who are not asked for severity will not distinguish, and the resulting list cannot be triaged. Reviewers who are asked generally provide it accurately, and the distinction is the single most useful thing a brief can add.

Where a reviewer marks many items as blocking, that is itself a signal worth investigating — either the output quality is genuinely poor or the reviewer's standard needs calibrating.

Asking for Specificity

Feedback that is actionable has a location and a proposed fix.

Ask reviewers to identify where each issue occurs — a timestamp, a segment number, or a quoted phrase — rather than describing it generally. "The terminology is inconsistent" is a starting point; "at 4:12 and 7:30 the same term is rendered differently" is a fix.

Ask for a suggested replacement where they have one. A reviewer who says a phrase reads as translated has identified a problem; one who supplies the natural alternative has solved it.

Ask them to note where they were uncertain. A reviewer's uncertainty is information, and it frequently points at genuine ambiguity in the source.

Calibrating Reviewers

Different reviewers apply different standards, and the variation is large enough to distort quality measurement.

Give the severity definitions with examples of each level. Abstract definitions are interpreted inconsistently; examples are not.

Explain what is out of scope. Reviewers frequently comment on the source content itself — its structure, its claims, its length — which may be useful but is not what was asked and should be routed separately.

Share previously approved assets as reference. A reviewer who can see what has been accepted before calibrates faster than one working from description.

Where multiple reviewers work on the same language, have them review the same asset occasionally and compare results. Divergence indicates the standard is not shared and needs clarifying.

For a new reviewer, review their first output alongside an established reviewer's, which surfaces calibration gaps immediately.

Turning Feedback Into Fixes

Feedback that is collected and not acted on trains reviewers to stop bothering.

Triage by severity. Fix blocking items, decide on quality items, record preference items without acting.

Route systematic findings to the source. A terminology issue that will recur belongs in the glossary, not just in this asset's fix list. A register problem that appears across assets belongs in the brief.

Close the loop. Tell reviewers what was applied and what was not, with brief reasoning for the latter. Reviewers whose feedback disappears into a queue disengage, and reviewer disengagement is one of the most common causes of stalled localization programs.

Track patterns. Which categories generate the most findings, in which languages, on which content types. This tells you where the process is weak, which is more valuable than any individual fix.

Update the glossary and the brief based on what recurs. A well-maintained brief gets shorter over time as settled questions move into the glossary.

Respecting Reviewer Time

Reviewers are the binding constraint in most programs, and using their time well is a practical necessity rather than a courtesy.

Send complete packages. A reviewer who has to request the source, the glossary, or the context loses time and momentum.

Batch requests rather than sending assets individually as they complete, particularly for reviewers in other time zones.

Be realistic about volume. Two to three hours per thirty minutes of content is a reasonable expectation for careful review, and asking for more in less time produces superficial review that provides false assurance.

Do not send content that has not passed mechanical checks. A reviewer catching encoding problems, timing faults, or missing segments is a reviewer whose expertise is being wasted on things a script should have caught.

Pay where you can. It changes the relationship, improves reliability, and makes it reasonable to expect turnaround.

A Template

A workable brief fits on one page:

The asset, its source, and the glossary. The audience, content type, and purpose. The register decision with reasoning. Known constraints. The specific questions, adapted to whether the deliverable is audio, subtitles, or both. The severity scale with definitions. The requested format: location, description, severity, and suggested fix. The deadline and the expected time commitment.

Reuse it across assets, adjusting only the asset-specific parts. Reviewers working from a consistent brief produce consistent output, which is what makes the feedback comparable across assets and over time.

The brief is a small artifact that determines whether the most expensive step in the localization process produces something useful. Most programs write it once and benefit from it indefinitely, and most programs that struggle with review quality have never written one at all.

Adapting the Brief by Content Type

The core questions hold across content, but the emphasis should shift with what is being reviewed.

Instructional content: prioritize whether a viewer could actually follow the instructions, whether terminology matches what they will see on screen, and whether pacing allows them to keep up. Comprehension matters more than elegance.

Marketing content: prioritize register intensity, whether claims land as intended rather than as overselling, and cultural fit. Ask whether the tone matches how brands in this market speak.

Regulated content: add a compliance dimension, and route it to a reviewer qualified in that domain rather than folding it into linguistic review. Ask specifically whether required language is present and correctly rendered.

Personality-driven content: ask whether the voice sounds like the same person and whether the humour and asides land. This is a different question from accuracy and requires a reviewer who is familiar with the original.

Technical content: ask whether a practitioner in the field would use this vocabulary, which is distinct from whether the terms are correct translations.

Children's content: add age-appropriateness as an explicit question, since complexity drift upward is the default failure and it is invisible to accuracy checking.

Adjusting the brief takes minutes and substantially improves the relevance of what comes back.

When Review Reveals a Process Problem

Individual findings are fixes. Patterns across findings are diagnostics, and reading them is where review earns most of its value.

Terminology findings concentrated in one subject area mean the glossary is incomplete there.

Register findings recurring across assets mean the register decision was not communicated, or was wrong.

Timing findings clustered at the ends of assets mean the timing process is not checking the end.

Pronunciation findings repeating the same terms mean pronunciation overrides are not persisting into the glossary.

Findings rising in one language mean something changed — a reviewer, a voice, a pipeline setting — and should trigger investigation rather than more review.

A high proportion of preference-level findings means the reviewer's standard needs calibrating, or the brief did not make clear what was in scope.

Feed each of these back into the process rather than fixing only the individual instances. A review programme that produces findings without changing anything upstream is generating the same findings indefinitely.

Reviewing at Scale

As volume grows, full review of every asset stops being possible, and the brief has to adapt.

Sampled review requires the same brief, applied to a randomly selected subset. The findings estimate the batch's quality rather than fixing every asset, and the sample must be genuinely random rather than convenient to be informative.

Targeted review narrows the questions rather than the assets. Asking every asset's reviewer only about terminology and factual accuracy, while reserving the full question set for sampled assets, covers the highest-severity categories broadly at lower cost.

Staged review puts different questions at different points. Terminology and accuracy at the script stage, where fixes are cheap; pronunciation and timing after generation, where they can only be assessed in the finished asset.

Automated pre-checks should run before human review so that reviewers never see assets with mechanical faults. A reviewer catching an encoding problem is expertise wasted.

Whatever the structure, keep the brief consistent across it. Findings collected under different briefs are not comparable, which defeats the purpose of tracking patterns over time.

Communicate the sampling honestly to stakeholders. A programme reporting that content is reviewed, when in fact a percentage is sampled, is creating a false impression of assurance that will eventually be discovered.