Recording studio with microphone and acoustic panels

A Decision Made Too Casually

Most localization decisions are invisible to the audience. Terminology, segmentation, and timing all matter enormously and none of them are consciously noticed when done well.

Voice is different. It is the first thing a viewer registers, it happens before comprehension, and it carries a judgement about the content's credibility and production quality that is formed within a second or two.

Despite that, voice selection is routinely left as a default — whatever the workflow picked, applied uniformly, with nobody making an explicit choice. The result is content where the voice is subtly wrong for the material in ways that depress engagement without anyone identifying the cause.

This is a short guide to making that decision deliberately.

The First Fork: Cloned or Assigned

Two fundamentally different approaches.

Voice preservation, or cloning, carries the original speaker's vocal identity into other languages. The viewer hears a version of the same person speaking their language.

Voice assignment uses a synthetic voice chosen for the target language, unrelated to the original speaker.

Cloning is right when the speaker's identity is part of the value. A creator whose audience follows them personally, a chief executive addressing employees, a well-known instructor, a recognisable brand spokesperson. In these cases the parasocial connection is a genuine asset and preserving it across languages retains something real.

Assignment is right when the speaker is functionally anonymous. Corporate narration, training content, help videos, product explainers, and most institutional communication. Here the viewer has no relationship with the original narrator, and a well-chosen native voice in the target language will usually outperform a cloned one — because a native voice can carry natural prosody and register in a way a cloned voice sometimes cannot.

There is a middle case worth naming. Where the original speaker is a professional narrator hired for the source language, cloning them into nine other languages preserves nothing the audience cares about while adding consent complexity. Assignment is clearly correct there, and it is frequently done the other way round by default.

Accent and Variety Within a Language

Having chosen a language, you still have to choose which version of it.

Most widely-spoken languages have several standard varieties, and the choice signals something.

Spanish for Spain and for Latin America differ noticeably, and Latin American Spanish itself spans varieties with distinct pronunciation and vocabulary. A neutral Latin American variety is the common default for pan-regional content.

Portuguese for Brazil and for Portugal are different enough that using one for the other audience is immediately noticed.

Arabic presents the sharpest version of this problem, with Modern Standard Arabic understood everywhere but sounding formal, and regional varieties sounding natural but limiting reach.

English across British, American, Australian, Indian, and other varieties, where the choice signals market focus.

French for France, Canada, and West Africa.

Chinese across Mandarin and Cantonese, which are different languages for these purposes, and across simplified and traditional script for text.

The practical guidance: choose deliberately based on where the audience actually is, use a neutral or pan-regional variety where you are serving several markets with one version, and be aware that a strongly marked regional accent will read as deliberately local rather than as neutral.

Voice actor recording in a professional booth

Matching Voice to Content

Beyond language and accent, the voice's character should suit what it is saying.

Considerations that recur:

Energy level. An enthusiastic product launch and a compliance training module need different delivery. A single default voice applied across a whole library will be wrong for a substantial part of it.

Authority versus warmth. Documentary, financial, and institutional content generally benefits from measured authority. Hospitality, healthcare, community, and consumer content generally benefits from warmth. Getting this backwards is a common and consequential error — a cold voice on patient education material undermines the content's purpose.

Age and gender. Should usually match the original speaker unless there is a specific reason otherwise, particularly in multi-speaker content where a mismatch confuses who is talking.

Pace. Instructional and safety content benefits from slower, more deliberate delivery than marketing content, and this should be a deliberate setting rather than a default.

Register consistency with the script. A formal translation delivered in a casual voice, or vice versa, produces a mismatch the viewer feels without identifying.

A useful test: play thirty seconds of the localized audio to a native speaker with no context and ask what kind of content they think it is. If their answer does not match what it actually is, the voice is wrong.

Consistency Across a Library

For anything beyond a handful of assets, consistency becomes the dominant concern.

One voice per language per content stream. A viewer working through a training series should hear the same voice throughout. Switching between modules is disorienting and reads as unfinished.

Maintain voice profiles for recurring speakers. A named presenter, host, or executive who appears repeatedly should have their voice choice recorded and reused, not re-decided per job.

Keep a documented voice map. Which voice is used for which language, which content stream, and which recurring speaker. Without this, a library localized over eighteen months by several people will drift.

Plan for continuity. If a voice becomes unavailable or a provider changes it, a library built on it needs a migration path. This is an argument for recording your voice decisions as data rather than leaving them implicit in each job.

Distinctness within multi-speaker content. Voices used for different speakers in the same asset must be easily distinguishable, particularly where speakers interrupt or exchange rapidly.

Cloning a real person's voice raises obligations that assignment does not.

Consent should be explicit, specific, and separate. Not bundled into a general content licence or employment agreement. It should state which languages, which content types, how long, and what happens on departure or withdrawal.

Control who can authorise content in that voice. An executive voice that anyone can generate content with is a governance gap. Define approval internally before the capability exists rather than after an incident.

Never synthesise a voice without the person's consent. This applies with particular force to interview subjects, sources, patients, students, and children, whose participation consent almost certainly did not contemplate synthesis.

Decide your disclosure position deliberately. Whether audiences are told a voice is synthetic is partly regulatory, partly contextual, and increasingly an expectation. Organisations that settle it in advance fare better than those asked unexpectedly. For journalism, healthcare, and public communication, disclosure is generally the lower-risk path.

Treat voice profiles as sensitive data. Several regulatory frameworks treat voice as biometric or biometric-adjacent. Isolate profiles per person and per tenant, control access, and support deletion.

Have a withdrawal path. If a person withdraws consent, you need to know which assets use their voice and be able to act.

Team discussing brand guidelines in a meeting

Reviewing Voice Output

Voice quality is not assessable by anyone who does not speak the language, which means review needs structure.

What a native reviewer should be checking:

Pronunciation of proper nouns. Names, places, products, and brands are the most likely errors and the most visible to viewers.

Tone and register match between the voice and the script's formality level.

Prosody and emphasis. Whether stress falls on the right words. Wrong emphasis can invert meaning while every word is correct.

Pace and naturalness, particularly whether the audio has been compressed to fit source timing.

Consistency across the asset, with no audible shifts mid-content.

Language-specific hazards. Tone accuracy in tonal languages, honorific level consistency where the language encodes it, gender agreement where the language requires it.

That last category is worth flagging to reviewers explicitly, because a reviewer checking for general naturalness may not think to verify honorific consistency across a twenty-minute piece.

A Working Checklist

  • Decide cloned versus assigned deliberately, based on whether the speaker's identity carries value.
  • Do not clone anonymous professional narrators; assign native voices instead.
  • Choose the language variety based on where the audience actually is.
  • Match voice energy, warmth, and pace to the content type rather than accepting a default.
  • Use one voice per language per content stream, and record the decision.
  • Maintain voice profiles for recurring speakers so choices persist across jobs.
  • Keep a documented voice map covering language, stream, and speaker.
  • Obtain explicit, specific, separate consent for any cloned voice, with a withdrawal path.
  • Define internally who may authorise content in a cloned voice.
  • Have native reviewers check proper nouns, prosody, register, and language-specific features such as tone or honorifics.

Auditioning and Standardising

For anything beyond a handful of assets, voice selection benefits from a short structured process rather than a per-project choice.

Audition against real content. Generate the same thirty-second passage from your actual material in several candidate voices. Abstract voice samples are a poor guide, because the fit between voice and content is what matters.

Include a native speaker in the decision. Non-speakers cannot assess naturalness, prosody, or register, which are precisely the criteria that separate candidates.

Test the hard cases. Proper nouns, technical terms, numbers, and any language-specific feature such as tone or honorifics. A voice that handles ordinary prose well may stumble exactly where your content lives.

Choose for the library, not the asset. The voice will carry a hundred videos, so evaluate it against the range of content you will produce rather than against the single piece in front of you.

Document the shortlist, not just the winner. When the chosen voice becomes unavailable, the runner-up is valuable information.

Re-audition periodically. Voice quality improves, and a library standardised three years ago may benefit from revisiting — balanced against the disruption of changing a voice audiences have grown used to.

Standardising once and recording the outcome removes a recurring decision, guarantees consistency, and makes any future migration a manageable task rather than a rebuild.

Frequently Asked Questions

Should I clone the original speaker or assign a native voice?

Clone when the speaker's identity carries value — a creator with a personal audience, a known executive, a recognisable spokesperson. Assign when the speaker is functionally anonymous, which covers most corporate, training, and help content. Cloning a hired professional narrator into nine languages preserves nothing the audience cares about while adding consent complexity, and assignment is clearly better there.

Which variety of a language should I choose?

Whichever matches where your audience actually is, chosen deliberately rather than by default. Where one version must serve several markets, use a neutral or pan-regional variety. Be aware that a strongly marked regional accent reads as a deliberate local signal rather than as neutral, which is right for regional content and odd for pan-regional content.

How do I keep voices consistent across a large library?

Record voice decisions as data rather than leaving them implicit in each job. Maintain a documented voice map covering which voice is used for which language, content stream, and recurring speaker, and reuse profiles rather than re-deciding. Libraries localized over many months by several people drift badly without this.

What consent is needed to clone someone's voice?

Explicit and specific consent, separate from any general content licence or employment agreement, stating the languages, content types, duration, and what happens if the person leaves or withdraws. Define internally who can authorise content in that voice, treat voice profiles as sensitive data with controlled access, and maintain a way to identify and act on assets using a withdrawn voice.

How can I review voice quality in a language I do not speak?

Structure the review. Ask a native speaker to check specific things: pronunciation of proper nouns, register match between voice and script, prosody and emphasis, pace and naturalness, consistency across the asset, and language-specific features such as tone accuracy in tonal languages or honorific consistency where the language encodes it. Unstructured requests to "check it sounds right" miss the systematic errors.


Related reading: Ethical Voice Cloning | Video Dubbing Voice Direction | AI Voice Cloning Explained