Language service provider team reviewing localized video

The Disintermediation Problem

Every language service provider has now had the conversation. A long-standing client asks why a thirty-minute training video costs what it costs, mentions that they ran a test through an AI dubbing tool over a weekend, and observes that the result was "basically fine."

The instinct is to explain why the result was not, in fact, fine. That instinct is usually correct on the merits and almost always wrong as a commercial strategy. The client is not asking for a defence of traditional dubbing. They are signalling that their reference price has moved, and that they expect their vendor to have an answer.

Agencies that treat this as a threat spend the next two years defending a shrinking book of business. Agencies that treat it as a service line design problem end up selling more, at better margins, to clients who now have budget for volumes they could never previously justify.

This piece is about the second path: what the service line actually looks like, how to price it, where the margin lives, and which parts of the traditional workflow survive contact with automation.

What Changes and What Does Not

It helps to be precise about which parts of a localization engagement are affected.

Substantially automated: transcription, first-pass translation, subtitle timing and segmentation, voice synthesis, format conversion, and batch processing of large libraries. These were once the bulk of billable hours on a video project. They are now closer to compute cost.

Largely unchanged: terminology governance, source-content consultation, in-market linguistic review, cultural adaptation, legal and regulatory sign-off, client relationship management, and the judgement about what should be localized at all.

Newly created: engine selection and configuration, glossary and custom-vocabulary maintenance across projects, prompt and context preparation, AI output triage, and the quality-assurance discipline needed to catch a different class of error than human translators produce.

The strategic reading is straightforward. The commoditized layer is the part that was always closest to commodity. The defensible layer is the part clients were never able to do themselves, and still cannot. The new layer is where early operators build advantage, because it requires accumulated process knowledge that a client running an ad-hoc weekend test does not have.

Pricing the Service Line

The single most common mistake is per-word or per-minute pricing carried over unchanged from the human workflow, discounted by some intuitive percentage.

This fails in both directions. On simple content, the discount is not aggressive enough and the client goes direct. On complex content, the discount is too aggressive and the agency absorbs review time it did not price for.

Three models work better in practice.

Tiered Output Quality

Sell three named tiers against explicit deliverables rather than against effort.

A raw tier delivers machine output with automated checks only, priced close to platform cost plus a modest handling margin. It suits internal-only content, archive material, and markets being tested rather than committed to.

A reviewed tier adds a single-pass linguistic review by an in-market reviewer against the source, with terminology enforcement and a documented error log. This is the volume tier, and it is where most enterprise content should land.

A certified tier adds a second reviewer, subject-matter validation, and a sign-off artefact. It suits regulated, safety-critical, and public-facing brand content.

Tiering works because it moves the conversation from "why does this cost more than the tool" to "which risk level does this asset warrant." Clients understand that framing intuitively, and they self-select upward more often than agencies expect.

Retainer on the Terminology Asset

The glossary, the custom vocabulary, the style guide, the approved rendering of every product name and regulatory phrase across every market — this is the asset that makes output good, and it compounds over time.

Charging a monthly retainer for maintaining it, separate from per-project fees, aligns incentives correctly. The client pays for the thing that actually determines quality. The agency is funded to maintain it between projects rather than rebuilding it under deadline pressure. And the asset becomes genuinely sticky in a way that a per-project relationship never is.

Volume Commitment with Banked Capacity

Enterprise clients with continuous content — support libraries, product releases, training refreshes — respond well to committed annual volume against banked processing capacity, with review services drawn down as needed.

This smooths agency revenue, gives the client a predictable unit economic, and surfaces the real planning conversation: how much content is there, in which languages, at what risk tier.

Team planning a multilingual content pipeline on a shared board

Where the Margin Actually Lives

Margin in an AI-assisted service line does not come from marking up processing. Processing is transparent, priced publicly by platforms, and trending toward zero.

Margin comes from four places.

Triage efficiency. The difference between a reviewer who watches every minute of output and a reviewer who knows which two minutes are likely to be wrong is a three-to-one cost difference on the same deliverable. That knowledge is process, and process is ownable.

Terminology reuse across clients in a vertical. An agency that has localized fifteen medical device training programmes has a working understanding of how regulatory phrasing must render across target markets. That accumulated judgement makes the sixteenth engagement faster and better than a generalist could deliver.

Source-side consultation. The highest-leverage intervention on most video localization projects happens before recording. Scripts written with localization in mind — avoiding idiom-dense phrasing, keeping on-screen text minimal, controlling speaking pace, avoiding culturally specific humour — produce dramatically better automated output. Selling this consultation is high-margin and structurally defensible.

Orchestration at scale. A client with four hundred videos across nine languages does not have a translation problem. They have a project management, versioning, approval-routing, and asset-delivery problem. That is agency work, and automation makes it larger rather than smaller, because the volume clients can now afford has grown.

Adapting Quality Assurance

The most important operational shift is that AI output fails differently than human translation, and existing QA processes are calibrated for the wrong failure modes.

Human translators produce errors that cluster around unfamiliarity: an unknown term, a misread of a technical concept, an inconsistency across a long document. Reviewers are trained to look for these.

Machine output produces errors that cluster around confidence. The system does not hesitate. It renders an ambiguous phrase fluently and incorrectly, and the fluency is precisely what makes the error hard to catch. A reviewer skimming for awkwardness will skim straight past a sentence that reads beautifully and says the wrong thing.

Practical adjustments that make the difference:

  • Review against source, never in isolation. Monolingual review of fluent machine output is close to useless. The reviewer must have the source in front of them.
  • Front-load terminology. Every error caught by glossary enforcement before generation is an error that never reaches a reviewer. Custom vocabulary is the cheapest quality intervention available.
  • Sample-and-escalate on low-risk tiers. Full review of every minute is not always warranted. Structured sampling with defined escalation thresholds is defensible and far cheaper.
  • Check the boundaries. Speaker transitions, numbers, dates, proper nouns, negations, and units are where automated output most often goes quietly wrong. A targeted checklist beats an unfocused full watch.
  • Log every correction. Errors that recur across projects are glossary entries waiting to be written. An error log that does not feed back into terminology is wasted effort.

Positioning Against the Client's Own Tools

Clients will run their own tools. This is fine, and fighting it is a losing position.

The productive posture is to be the party that tells the client honestly which content should go through their own tool untouched, which should come to you reviewed, and which should never be automated at all. An agency that recommends against its own reviewed tier on low-stakes internal content earns the credibility to insist on the certified tier when a product safety video is on the table.

Concretely, three positions hold up well.

Be the risk assessor. Most clients cannot articulate which of their assets carry regulatory, brand, or safety exposure in which markets. An agency that can produce that map is selling judgement, not throughput.

Be the consistency layer. A client running their own tool ad hoc across six departments will produce six different renderings of their own product names. Centralized terminology governance is unglamorous and enormously valuable.

Be the escalation path. Clients hit hard cases — a language pair that performs poorly, content with heavy accent variation, a market with regulatory requirements they did not anticipate. Being the reliable answer to "this one didn't work" is a durable position.

Reviewer comparing source and translated transcript side by side

Building the Operational Foundation

A few concrete requirements separate agencies that scale this cleanly from those that struggle.

Workspace separation per client. Client assets, glossaries, and voice profiles must not bleed across accounts. Platform-level workspaces with role-based access are the baseline, not a nice-to-have, and clients in regulated sectors will ask about it during procurement.

Confidentiality posture in writing. Clients will ask whether their content trains a model, where it is processed, and how long it is retained. Having a clear, accurate, documented answer — including the platform's position on not training on customer data — removes a common procurement blocker.

API-driven intake for volume accounts. Manual upload works to a point. Past that point, programmatic submission, status polling, and automated delivery into the client's asset system is what makes a four-hundred-asset engagement profitable.

Reviewer network by vertical, not just by language. A fluent reviewer without domain knowledge will pass a technically wrong sentence. For regulated verticals, domain competence matters more than the last increment of linguistic polish.

A Realistic First Ninety Days

Agencies that launch this well tend to follow a similar sequence.

Weeks one to three: pick two existing clients with recurring video volume and offer a no-charge pilot on a representative asset in two languages. Use it to calibrate internal review time, not to impress the client.

Weeks four to six: build the triage checklist and the error log from what the pilot surfaced. Define the three tiers against your actual measured effort rather than against a guess.

Weeks seven to ten: price the tiers, write the terminology retainer into a proposal, and take it to the two pilot clients. Expect the conversation to be about risk, not about rate.

Weeks eleven to thirteen: instrument everything. Cost per finished minute by tier, review hours per source hour by language pair, error categories by frequency. Without this, the second year of the service line is priced on instinct.

Frequently Asked Questions

Will offering AI dubbing cannibalize our existing human dubbing revenue?

Partially, and the alternative is worse. The revenue most at risk is the high-volume, low-complexity work that clients will automate with or without you. Agencies that add the service line typically see per-project revenue fall and project count rise, with total revenue growing because clients localize content they previously could not justify. The engagements that stay human — brand campaigns, performance-critical narration, regulated markets requiring certified human translation — are also the higher-margin ones.

How do we handle clients who ask why they should pay us at all?

Answer the question directly rather than defensively. They should pay for terminology governance, in-market review, risk assessment, and orchestration. They should not pay a large markup on processing. Being explicit about which is which builds more trust than a bundled rate that invites suspicion.

What margin should we expect on a reviewed tier?

It varies by language pair and content complexity, but the dominant variable is review hours per source hour, which falls significantly as your triage process matures. Agencies commonly find that first-quarter review effort on a new client is double what it settles at by the third quarter, which is why pricing on early-engagement effort leaves margin on the table.

Do we need to disclose to clients that output is machine-generated?

Yes, and it should be in the statement of work rather than buried. Clients in regulated sectors may have obligations around it, and discovering it later damages trust disproportionately. Framing tiers explicitly around the degree of human involvement makes the disclosure a feature of the offer rather than an awkward footnote.

Which content should we advise a client not to automate?

Performance-driven narration where delivery is the product, content in language pairs where your reviewers consistently report high correction rates, material with heavy overlapping speech or strong dialect variation, and anything where a translation error creates legal, safety, or clinical exposure that a review pass cannot fully retire.


Related reading: Enterprise Video Localization | Localization Quality Assurance Checklist | Video Translation Pricing Models