Video translation pricing and cost models

Why Comparison Is Hard

Requesting quotes for video translation from several providers typically produces numbers that cannot be compared directly, because they are pricing different scopes.

One quote covers transcription, translation, and a subtitle file. Another covers the same plus generated audio. A third covers everything plus native-speaker review and on-screen text replacement. The headline per-minute rates might be within a few percent of each other while the actual deliverables differ by a factor of three in labor.

Understanding the pricing models, and what sits inside and outside each, is the only way to make a real comparison. This guide covers the common structures, what drives cost within each, and the expenses that are frequently excluded from the quoted figure.

Per-Minute Pricing

The dominant model prices by minute of source video, per target language.

How it works. A rate is quoted per minute per language. A thirty-minute video into four languages at a given per-minute rate costs thirty times four times that rate. Some providers round up to whole minutes; some bill in fractional increments.

What drives the rate. Scope is the largest factor — subtitles only versus subtitles plus audio versus a fully finished localized video. Language pair matters, with common pairs cheaper than rare ones. Turnaround time affects rate, with expedited work priced higher. Content complexity matters, with technical and specialized content commanding higher rates than general content.

Strengths. Predictable, easy to budget, and scales cleanly. For irregular or project-based needs, it avoids paying for capacity you do not use. Good for one-off projects and for organizations with unpredictable volume.

Weaknesses. Expensive at volume compared with subscription models. Creates a per-item decision that discourages localizing marginal content — every additional video is a fresh spending decision, which tends to suppress volume even when the aggregate return would justify it.

What to check. Whether the rate is per minute of source or per minute of output, whether partial minutes round up, whether revisions are included or billed separately, and precisely which deliverables the rate covers.

Subscription and Credit Models

Subscription pricing provides a monthly allowance of processing minutes for a fixed fee.

How it works. A monthly or annual fee includes a defined number of minutes. Additional minutes are either unavailable, purchasable at an overage rate, or accessible on a higher tier. Some models use credits that can be spent across different operations rather than a straight minute allowance.

Strengths. Substantially lower effective per-minute cost at consistent volume. Removes the per-item spending decision, which in practice increases how much gets localized — content that would not survive an individual cost justification gets processed because the capacity is already paid for. Predictable budgeting.

Weaknesses. Requires consistent volume to be efficient. Unused allowance is typically wasted, so lumpy demand fits poorly. Overage rates can be unfavorable.

What to check. Whether unused minutes roll over, how overage is priced, whether minutes are consumed per language or per source video — this distinction matters enormously, since a thirty-minute video into six languages consumes either thirty or one hundred eighty minutes depending on the model. Also check whether reprocessing after a correction consumes the allowance again.

Where it fits. Organizations with steady ongoing localization needs — content programs publishing regularly, product teams maintaining a library, creators localizing continuously. Octavia's plans follow this structure, which suits recurring localization rather than one-off projects.

Enterprise and Contract Pricing

Larger deployments typically move to negotiated arrangements.

How it works. Volume commitments in exchange for reduced rates, often with committed annual spend, dedicated support, custom terms, and integration or API access included. Terms vary widely and are genuinely negotiable.

What is typically included beyond rate. Service level commitments on turnaround, data processing terms and residency guarantees, security review accommodation, dedicated account management, custom terminology and voice configuration, and integration support.

Strengths. Best effective rate at high volume. Contractual commitments on turnaround and data handling that self-serve tiers do not provide. Procurement and compliance requirements can be addressed properly.

Weaknesses. Commitment risk if volume forecasts prove wrong. Longer sales and procurement cycle. Less flexibility to switch providers.

What to check. What happens to unused commitment, how rates adjust if volume exceeds forecast, whether the service level commitments carry meaningful remedies, and what the data processing terms actually specify — particularly processing location, retention, and whether content is used for model training.

Human Translation Pricing

Where human translation is required — for certification, regulated content, or quality levels that automated processing does not reach — pricing follows different conventions.

Per word is standard for translation of transcripts and scripts. Rates vary by language pair, subject matter, and translator qualification. Certified and sworn translation commands higher rates.

Per hour is common for review, editing of machine output, and subtitling work where the task is not cleanly measurable in words.

Per minute of video appears for subtitling and captioning services that include timing work.

Minimum charges apply widely and matter for short content. A two-minute video may cost the same as a ten-minute one under a minimum.

Post-editing of machine translation is usually priced between raw machine translation and full human translation, and the rate depends on the quality level required — light post-editing for comprehensibility costs considerably less than full post-editing to publication standard.

What Drives Cost

Independent of model, certain factors move cost predictably.

Number of languages is the largest multiplier and scales close to linearly for human-involved stages while scaling more cheaply for automated ones.

Deliverable scope. Subtitles only is the cheapest tier. Adding generated audio increases cost modestly in automated workflows. Adding native-speaker review increases it substantially. Adding on-screen text replacement and re-recorded screen captures increases it a great deal, because these are production work rather than language work.

Language pair. Common pairs have more capacity and lower rates. Low-resource languages cost more and may have limited availability.

Subject complexity. Technical, medical, legal, and financial content requires reviewers with domain knowledge, who cost more and are harder to schedule.

Source quality. Poor audio increases transcription correction time. Unclear speech, heavy accents, background noise, and overlapping speakers all add labor.

Turnaround. Expedited work carries premiums, sometimes substantial.

Revision cycles. Whether revisions are included, and how many, materially affects total cost on content that goes through stakeholder review.

Costs Frequently Excluded

The gap between a quote and the true program cost is usually made up of these.

Internal review time. Native-speaker review by your own staff is not free even when it is not invoiced. For a thirty-minute video, budget two to three hours of reviewer time per language. Across a program this is often the largest real cost and it appears in nobody's quote.

Project coordination. Submitting files, tracking status, chasing reviewers, and publishing outputs is a real workload at volume, frequently amounting to a full role.

On-screen text and graphics. Almost never included in a translation quote. This is design and production work, and for graphics-heavy content it can exceed the translation cost.

Screen re-recording for software content, along with maintaining localized demo environments.

Metadata localization. Titles, descriptions, tags, and thumbnails need localizing for the content to be discoverable, and this is usually outside scope.

Maintenance. Translated content goes stale when the source changes. Ongoing update cost is rarely modeled at program inception and frequently exceeds the initial translation cost over a few years.

Storage and delivery. Multiple language versions multiply storage and bandwidth.

Rework. Errors found after publication cost more than the original processing, particularly for burned-in output requiring re-render.

Comparing Quotes Accurately

Build a specification before requesting quotes, and require quotes against that specification rather than accepting each provider's default scope.

The specification should state: source video count and total duration, target languages, required deliverables per language, whether native-speaker review is included and by whom, revision allowance, turnaround requirement, file formats, and data handling requirements.

Then compare total program cost rather than headline rate: quoted cost plus estimated internal review hours plus coordination plus excluded production work plus expected maintenance over the content's life.

Request a paid pilot on real content before committing at volume. A single representative video processed by each candidate provider, reviewed by your own native speaker, reveals quality differences that no specification captures. The cost of the pilot is trivially small against the cost of discovering quality problems after committing a year's volume.

Check the data terms in writing, particularly for regulated or confidential content. Verbal assurances about processing location, retention, and training use are not sufficient where compliance obligations apply.

Building a Realistic Budget

A defensible budget models the full program rather than the processing line item.

Start with volume: total source minutes to be localized in the period, and the number of target languages. Multiply for the languages where costs scale per language, which is most of them.

Add processing cost using whichever model applies. This is usually the smallest component in a program with meaningful review requirements.

Add internal review hours at a realistic rate. Two to three hours per thirty minutes of content per language is a workable planning figure for general content, higher for technical or regulated material. Cost it at the loaded rate of whoever performs it, whether or not it appears in a budget line.

Add coordination. At low volume this is absorbed into someone's existing role; above roughly fifty assets per period it becomes a distinct workload.

Add production work: on-screen text, graphics, screen re-recording, and metadata localization. For graphics-heavy or software content this is frequently the largest single component and is almost never in the quote.

Add maintenance as a recurring annual figure. A reasonable planning assumption is that a meaningful proportion of the library will need updating each year, scaling with how fast the underlying subject changes.

Add a rework allowance. Errors will be found after publication, and the correction cost exceeds the original processing cost.

The resulting figure will exceed any vendor quote substantially. That is the point — it is the number that determines whether the program is viable, and discovering it after committing is how localization programs stall halfway through a catalog.

Where Costs Have Actually Fallen

It is worth being precise about which parts of video localization have become cheap, because the savings are uneven and budgeting on an average is misleading.

Transcription, translation, voice generation, and subtitle production have fallen dramatically in cost. Work that required studio time, voice talent, and specialist labor now runs on software at a small fraction of the previous cost, and the quality gap has narrowed considerably.

Native-speaker review has not fallen. It requires a qualified person's attention for a duration proportional to the content, and no tooling changes that arithmetic.

Production work — graphics, on-screen text, screen recording — has not fallen meaningfully either. It remains design and production labor.

Coordination has fallen only where it has been automated. Manual file handling costs the same as it always did.

The practical implication is that the cost profile of a localization program has inverted. Where processing once dominated the budget and review was a modest addition, review and production now dominate while processing is a minor line. Programs budgeted on the old proportions consistently underestimate, because they scaled the wrong component.

Choosing a Model

Occasional projects with unpredictable timing: per-minute pricing, avoiding commitment to capacity you will not use.

Steady ongoing volume: subscription, which lowers effective cost and removes the per-item decision that otherwise suppresses volume.

High volume with compliance requirements: enterprise contracting, for the rate, the service levels, and the data terms.

Regulated, certified, or evidentiary content: human translation pricing for the portions requiring it, with automated processing for triage and drafting.

Most mature programs use more than one. A subscription covers the routine ongoing work, per-minute or human pricing covers the exceptions, and the mix shifts as volume and requirements change.

The most common budgeting error is comparing quotes on the automated processing cost, which is the part that has become cheap, while omitting the review and production costs, which are the parts that have not. Modeling the full cost — including your own team's hours — produces a number that is higher than any quote and considerably more accurate.