A dubbing project rarely fails its first review because the translation was wrong. It fails because the voice was wrong, and nobody in the room can say precisely why. The lines are accurate, the timing is close, the audio is clean, and the client still comes back with "not quite right." That note is impossible to act on, so the team re-records, guesses again, and burns a cycle. Two or three cycles later, the launch date is the only thing still moving.

The missing artifact is a voice casting brief. It is a short document, usually one to three pages, that settles the decisions a voice has to satisfy before anyone auditions. It names the role, the audience, the register, the age range, the energy, the pace, and the qualities that separate an acceptable read from the correct one. Its purpose is not to be elegant. It is to make a casting director in one country, a reviewer in another, and a brand manager in a third judge the same thing.

The brief matters more in dubbing than in original production because the people applying it sit in different time zones and often do not share a native language. A note like "warmer" survives a desk conversation and dies in a translated email. A brief that says "lets the end of a sentence fall instead of lifting, and holds volume steady when the script gets excited" survives both.

What follows covers what a brief contains, how to describe voice quality so a reviewer in another language can apply it, how to hold brand voice across a set of different speakers, when cloning beats casting, and how to document the decision so episode twelve inherits the reasoning from episode one.

The fields a voice casting brief must define

Most briefs arrive as one sentence: "warm, friendly, professional, male, thirties." Every word in it means something different to each reader, and two casting directors will return two different slates that both technically satisfy it. The fix is not better adjectives. It is a fixed field set with written values.

The minimum set:

  • Role and function: whether the voice is on camera, off camera, narrating, or a character in dialogue, and whether it guides, explains, questions, or opposes.
  • Audience: who is watching, in what setting, and what they already know about the subject.
  • Market: the language and territory the voice is cast for, and the delivery conventions that audience expects.
  • Register: the formality of the delivery, which is not the formality of the script.
  • Age band: a range such as 30 to 45, never a single number.
  • Energy: sustained, conversational, or deliberately restrained, stated as a target rather than a mood.
  • Pace: a words-per-minute band with a tolerance at each end.
  • Warmth: whether the voice leans toward the listener or presents from a distance.
  • Texture: breathiness, rasp, brightness, resonance, and anything that would disqualify a candidate no matter how well they read.

Turning field values into something a candidate can hit

A field value only a linguist can apply is unfinished. Warmth becomes usable when it reads as "addresses one listener closely, keeps volume even when the script gets excited." Pace becomes usable when it reads as "155 to 175 words per minute, may drop to 145 for technical passages, never exceeds 185." Each field should describe something a candidate can choose to do on the next take.

What the brief should not try to control

Casting is not script editing, mixing, or translation. A line that needs a rewrite to land in a target language is a translation decision, and folding it into the brief confuses candidates about what they are being asked to change.

The one-page test

If a candidate's agent can read the brief and predict the three things that will get their client rejected, the brief is finished. If they cannot, it is still a description of a feeling.

Describing voice qualities so a reviewer in another language can apply them

Adjectives describe the effect on the listener, not the behavior of the speaker. Two reviewers in different languages map "warm" onto different behaviors: one hears lower volume and slower pace, the other hears a smile in the voice. When they disagree about a candidate, they are usually disagreeing about the definition rather than the audio.

Replace adjectives with observable behavior

Behavioral descriptors give both reviewers the same thing to listen for:

  • Sentence-final behavior: does the pitch fall, stay level, or lift?
  • Volume stability: does loudness track the emotional cues in the script, or stay even through them?
  • Distance: does the voice address one listener closely, or present to a room?
  • Pause behavior: does the speaker pause for punctuation, or for thought?
  • Consonant energy: crisp and forward, or soft and blended?

Anchor every descriptor to a reference and a counterexample

Give one clip that is right and one that is close but wrong, with a note on the single dimension that fails. This teaches the definition faster than a word list and gives the reviewer something to point at when they argue.

Keep a shared vocabulary sheet

Build a one-page glossary with a timecode example for each term, then translate the glossary, not just the brief. A reviewer applying notes in Japanese, Polish, or Portuguese needs the same definitions the original casting director used, or the two are running different projects that happen to share a script.

Separate performance from production

"Recorded in a small room" is a technical fault, not a casting note. A brief that mixes the two invites candidates to solve a problem they cannot solve from a booth in another country.

Keeping brand voice consistent when every target voice is a different person

A brand does not have one voice. It has one set of behaviors that different speakers in different languages approximate. Consistency comes from deciding which behaviors are non-negotiable and which are expected to shift. On-screen text carries register decisions of its own, so the same logic applies when you handle subtitle translation alongside the dub.

Split the brief into an invariant layer and a local layer

The invariant layer holds pace band, energy curve across a paragraph, sentence-final behavior, distance from the listener, how humor is handled, and how the brand name is pronounced. The local layer holds perceived age, perceived authority, formality register, forms of address, and dialect expectations. Mark each field as one or the other. Unmarked fields are where reviewers argue.

Adapt the traits, not the adjectives

"Young and current" resolves differently in different markets: bright and fast in one, low and dry in another. If the brief specifies the adjective, the target casting director hands back a local cliché of the adjective. If it specifies the behavior, they cast for the behavior.

Judge consistency across the set, not language by language

Play the first thirty seconds of each language back to back and ask whether a listener who understands none of the languages would describe them as one brand. This is also how you catch the language that sounds correct in isolation and wrong in company.

Using reference clips without asking for an impersonation

What the reference clip is for

A reference clip is a calibration instrument. It demonstrates pace, energy, sentence-final behavior, and distance. It does not specify an identity. Candidates who treat it as a template produce a shallow copy; candidates who treat it as a measuring stick produce something that fits the role.

Label every reference with the dimension it illustrates

Write the timecode, what to listen for, and what to ignore, in that order. A usable label reads: "Reference A, 0:14 to 0:22, sentence-final fall and even volume. Ignore the accent, the recording quality, and the subject matter." Without the ignore clause, every candidate copies the accent.

The impersonation trap

Never write "we want someone who sounds like" a named person. That instruction asks for a surface rather than a fit, it demands a performance few candidates can sustain across a full script, and in many jurisdictions a recognizable imitation of a living person carries right-of-publicity exposure. Ask for the pacing of one voice and the warmth of another, and state plainly which aspects are not wanted.

Keep some references internal

A rejected voice from a previous round is useful calibration for the casting team and harmful to candidates. Keep a short internal appendix for clips the candidate should never see.

When to clone a speaker instead of casting

Casting assumes the voice is interchangeable. Sometimes it is not, and preserving the original speaker across languages is the correct answer rather than a shortcut.

The cases where cloning is the right call

  • The speaker is the brand asset: a founder, a named expert, a presenter whose identity is why people watch.
  • Series continuity: the voice itself is the recognition cue across episodes or seasons.
  • Identity-carrying content: testimony, executive announcements, statements where the audience must know who said it.
  • Volume: the speaker cannot re-record for every language and every content update.

Authorization should be written and specific: which languages, which content types, what term, whether new scripts may be generated, and who signs off. Disclosure is a policy decision before it is a legal one, and several jurisdictions require that synthetic voice use be disclosed. Decide where the disclosure lives, whether that is the description, an end card, or an on-screen label, and apply it in every language.

What cloning does not fix

Cloning preserves timbre and delivery habits. It does not translate, direct, or adjust pace. A clone of a fast speaker is still fast in a language that needs more syllables. The performance still needs a director, and the same speech generation pipeline that produces the dub still needs a brief. Where the per-language cost of a full cast against cloned output decides the approach, the plans are laid out on the pricing page.

Multi-speaker content and the voice map

Building the map

A voice map is a table listing every speaking role, its brief identifier, the cast voice per language, and the episodes it covers. Build it before the first recording session rather than after the first consistency complaint. Add a register distance column: two roles with the same age band, energy, and register are hard to tell apart, so the map has to record how they differ.

Keeping the map stable across episodes

Listeners notice a voice that shifts slightly between episode two and episode three more than they notice an imperfect but constant voice. Lock the cast per role per language for the life of the series. If a change is unavoidable, make it at a season boundary rather than mid-arc.

Overlapping speech and speaker separation

Multi-speaker source audio often arrives with people talking over each other. Speaker diarization establishes who spoke when, and dialogue and music separation isolates speech before the audio translation pass begins. Without those steps, a casting decision cannot be applied cleanly to the right line, and the dub will attribute the wrong voice to the wrong character.

Roles that return after a long gap

If a voice is unavailable for a returning character, note the reason in the map. Otherwise the next person re-casts from memory and the character changes without anyone deciding to change it.

Shortlisting candidates on the same test lines

Choosing test lines

Three to five lines stress different dimensions at once:

  • A long expository sentence, to test breath control and pace.
  • A line containing a proper noun, a number, and a unit, to test pronunciation.
  • An emotional beat, to test range without pushing the candidate into performance.
  • A line with humor, to test timing.
  • The brand's own tagline, read the way it appears in the content.

The comparison grid

Score each candidate on a fixed set: pace, energy, sentence-final behavior, distance, pronunciation of brand and product names, and overall fit. A number alone is not enough. Each score needs one sentence of justification, because that sentence is what makes the decision reviewable three months later.

Listening conditions

Same headphones, same loudness normalization, and the candidate's name hidden where possible. A few decibels of loudness difference changes perceived warmth and confidence, which is enough to swing a close call between two competent reads.

The shortlist process, in order

  1. Send the brief, the labeled references, and the test lines to casters in each target language.
  2. Ask for three candidates per role, each reading the identical lines.
  3. Normalize loudness across all submissions before comparing anything.
  4. Blind-listen to the full set once without scoring, then a second time with the grid open.
  5. Shortlist two candidates per role and record them again on the full emotional range the role requires.
  6. Decide, write the reason, and archive the losing takes beside the winning one.

Running the comparison inside a video dubbing workflow keeps the test lines, the language, and the role versioned in one place, which matters once the project passes its tenth deliverable.

Documenting the decision and the reusable voice casting brief template

The casting record

Append one page to the brief. It carries the date, the decision-maker, the candidates reviewed, the chosen voice, the reason in one or two sentences, and the file names of the winning takes. When someone asks in episode nine why the narrator sounds the way it does, the answer is on that page rather than in someone's memory.

Change control

Any change to a field after casting begins should be dated and attributed. A brief that quietly shifts between episode three and episode eleven is how a series acquires a voice that matches nothing, and no single person notices until a viewer does.

The reusable voice casting brief template

The fields that matter, in order:

  • Project name and content type
  • Language and territory
  • Role name and function in the content
  • Audience and viewing context
  • Register
  • Age band
  • Energy target
  • Pace band in words per minute, with tolerance at each end
  • Warmth and distance
  • Texture notes and disqualifiers
  • Pronunciation rules for the brand name, product names, numbers, and units
  • Reference clips, each labeled with the dimension it illustrates
  • Counterexample clips, each labeled with the dimension that fails
  • Test lines
  • Delivery requirements: file format, sample rate, channel count, noise floor
  • Approval chain and turnaround
  • Consent and disclosure status, where the voice is cloned

Teams that assign voices in software can hold these fields as structured data attached to a role, so the brief travels with the assignment instead of living in a separate document that goes stale. The API documentation shows how structured metadata and voice identifiers are attached at the project level.

Frequently asked questions

Does the whole brief need translating into every target language?

No, but the vocabulary sheet does. Reviewers working from their own translation of "warm" are not applying the same instruction, and the disagreement surfaces as inconsistent casting rather than as a translation problem.

How long should a voice casting brief be?

One to three pages for a single role. Longer briefs usually mean translation notes, script notes, or mixing notes have been folded into casting, and each addition makes the actual voice criteria harder to find.

Can one voice cover more than one role?

Sometimes, if the roles never appear in the same scene and the register distance is wide enough. Record the overlap in the voice map so no listener hears the same speaker on both sides of a conversation.

What if the client rejects every candidate?

That usually means the brief describes an effect no speaker can produce, or two fields are in conflict, such as a sustained high-energy target combined with a deliberately restrained register. Re-check the field set before running another round.

Is a reference clip the same as a style guide?

No. A style guide describes the brand's language and visual rules; a reference clip illustrates one dimension of a voice. Keeping them separate prevents a candidate from being asked to imitate a look instead of a delivery.

When is cloning the wrong choice?

When the role is a character rather than the speaker, when the speaker has not given written authorization covering the intended languages and content types, or when the audience would reasonably feel misled by not knowing the voice is synthetic.

Conclusion

The next time a voice comes back wrong, the first question is whether a brief existed. In most cases it did not, and the fix is not a new round of candidates. It is a page of decisions: the fields, the behavioral definitions, the labeled references, and the scored test lines, filed somewhere the next episode can find them.

Start with one role on the next project rather than the whole cast. Write the invariant layer, pick three test lines, run three candidates per language, and score them on the same grid. That produces a decision you can defend to a client and reuse next season without rebuilding the reasoning from scratch.

Where the voice has to be the original speaker, treat cloning as a casting decision with its own consent and disclosure requirements rather than a technical convenience. Where the content is episodic and multi-speaker, let the voice map do the remembering. Both belong in the same file as the brief, which means the next project begins with a document instead of a guess.