The Problem With Most Localization RFPs
A typical video localization RFP asks how many languages a vendor supports, what their turnaround time is, whether they have quality processes, what their pricing is, and whether they can handle enterprise volume.
Every vendor answers yes, many, fast, and yes. The responses are indistinguishable, the evaluation defaults to price and presentation, and the differences that actually determine whether the engagement works — how change is handled, who owns the assets, what happens when quality is disputed — surface after signature.
A good RFP is built from questions where a capable vendor's honest answer differs from a weak vendor's honest answer. That means asking about specifics, mechanisms, and edge cases rather than capabilities.
The other function of a good RFP is to force your own side to decide things. A large share of failed localization engagements fail because the buyer had not determined what quality meant, who would review, or what the content pipeline looked like — and no vendor can supply those.
Define Your Scope Precisely
Vague scope produces vague pricing and disputes later. Be specific about:
Content volume and shape. Hours per month, average asset length, and the distribution. Two hundred hours of ten-minute training modules is a completely different engagement from two hundred hours of ninety-minute webinars, and vendors price them differently.
Content types. Scripted narration, presenter-to-camera, panel discussion, screen recording, dramatised scenario, live event capture. Difficulty and cost vary enormously and panel content in particular is dramatically harder than scripted narration.
Source language and target languages, with the regional variants specified. "Spanish" is not a specification; European or Latin American Spanish is, and the difference matters commercially.
Deliverables per language. Subtitle tracks, dubbed audio, SDH captions, transcripts, audio description, localized on-screen graphics, localized metadata. Each is a separate line and buyers routinely omit some and then expect them.
Source material availability. Whether you can supply project files, separate audio stems, multitrack recordings, scripts, and existing glossaries. This affects both cost and achievable quality more than almost anything else, and vendors will assume the worst if you do not say.
Revision expectations. How often content is re-cut, and what the expected turnaround on a revision is.
Volume trajectory. Whether this is steady state or growing, and how fast.
Questions That Separate Vendors
"Walk us through what happens when we re-cut a video after delivery." The best question on this list. A capable vendor describes segment-level diffing, regeneration of only affected portions, and preservation of reviewed content elsewhere. A weak one describes re-doing the asset and charging again. This single question predicts the ongoing cost of the relationship better than the rate card.
"How do you handle a term we want rendered differently after fifty videos are delivered?" Tests whether terminology is managed as a durable asset applied retroactively or as per-job instructions. Ask specifically whether previously delivered assets are updated and at what cost.
"Who owns the translation memory, glossaries, and voice assets at the end of the contract, and in what format do we receive them?" The most consequential commercial question and the one buyers most often forget. Assets you cannot take with you are a lock-in mechanism. Get the answer in the contract, with a specified format and a delivery timeline.
"Show us a sample in a language we can have independently assessed." Not a showreel — a sample of the kind of content you actually have, in a language where you can commission an independent review. Vendor showreels are their best work on their easiest content.
"What happens when we reject a delivery?" Tests whether there is a defined remediation process, a turnaround commitment, and a threshold at which rejection has commercial consequences. Vague answers here mean disputes will be resolved by whoever is more persistent.
"How do you measure quality, and what will you report to us?" Look for specific, measurable things — reading speed compliance, terminology adherence rate, error categories and counts, review pass rates — rather than a description of a process. Ask what the report looks like and how often it arrives.
"Which parts of your process are automated and which are human?" Every vendor uses automation now. Evasiveness about it is a signal in itself; a clear answer lets you judge whether the human effort is placed where it matters. There is nothing wrong with a heavily automated pipeline as long as review is applied where the risk is.
"How do you handle content we cannot send outside our environment?" Relevant for regulated, pre-release, or confidential material. Ask about data residency, retention, subprocessors, and whether the material is used for model training.
Pricing Structures and What They Hide
Compare structures, not just numbers, because different units make quoted rates incomparable.
Per source minute is the most common and the easiest to compare. Confirm whether it is per source minute per language, and whether it varies by content type or language tier.
Per word is inherited from document translation and fits video badly, because it ignores the timing, mixing, and rendering work that dominates the effort.
Per output asset can look attractive and hides variability in asset length.
Subscription or committed volume trades flexibility for rate, and the question to ask is what happens to unused volume and what the overage rate is.
Costs that frequently sit outside the headline rate:
- Setup and onboarding per language
- Voice creation or cloning
- Terminology extraction and glossary build
- Review cycles beyond an included number
- Rush and expedited delivery
- Re-cut and revision handling
- Format conversion and platform-specific packaging
- On-screen text and graphics localization
- Asset storage and retrieval after a retention period
- Asset export at contract end
Ask for a worked example: a specific, representative asset priced end to end into three languages with everything included. This surfaces the excluded items faster than reading a rate card.
Evaluating the Responses
Weight the pilot heavily. A paid pilot on your real content, evaluated by your reviewers, is worth more than the entire written response. Structure the RFP so the shortlist runs one, with the same brief and the same content for every vendor.
Use your own reviewers, blind. Native speakers who know your subject matter, assessing unlabelled outputs against a defined scorecard. Vendor-supplied quality assessments assess the vendor.
Score the operational answers, not the capability claims. Re-cut handling, terminology retroactivity, asset ownership, and rejection process are where engagements succeed or fail.
Check references on the specific things you care about. Not "were you happy" but "what happened when you re-cut a video", "how long did it take to fix a terminology decision", "what did the first six months cost against the estimate".
Model three-year total cost, not year one. Setup costs are one-off, volume grows, and exit costs are real. A vendor cheaper in year one can be substantially more expensive across the term.
Test the escalation path. Ask who you call when something is wrong at 6pm before a launch, and whether that is contractual or aspirational.
Terms Worth Getting Right
Asset portability. Translation memories, glossaries, voice models where licensing permits, reviewed transcripts, and timing data — delivered in a documented, non-proprietary format within a specified period after termination, without additional fee.
Data handling. Residency, retention, deletion on request, subprocessor disclosure, and an explicit position on whether your content is used to train models.
Voice rights. If voices are cloned from your people, the contract must state that you control them, that the vendor cannot use them elsewhere, and what happens to them at termination.
Quality thresholds with consequences. A defined standard, a measurement method, and a remedy when it is not met. Quality language without a remedy is decoration.
Change control. How rates change, with how much notice, and what happens when your volume moves substantially in either direction.
Transition assistance. An obligation to support migration to another provider, specified in the contract rather than negotiated at the point you have decided to leave.
Confidentiality that covers pre-release content specifically, if you handle it.
A Working Checklist
- Specify volume, asset length distribution, and content types rather than hours alone.
- Name regional variants, not just languages.
- List every deliverable per language explicitly.
- State what source material you can supply, including stems and project files.
- Give the expected revision frequency and trajectory of volume.
- Ask how re-cuts are handled and priced.
- Ask whether terminology changes are applied retroactively and at what cost.
- Require asset ownership, export format, and export timeline in writing.
- Request a sample on your own content, not a showreel.
- Ask for the rejection and remediation process with turnaround commitments.
- Ask which steps are automated and which are human.
- Ask for data residency, retention, subprocessors, and model training position.
- Compare pricing structures and request a worked end-to-end example.
- Enumerate the costs that typically sit outside the headline rate.
- Run a paid pilot with identical content across the shortlist.
- Evaluate blind, with your own native reviewers, against a defined scorecard.
- Check references on specific operational events rather than general satisfaction.
- Model three-year total cost including setup and exit.
- Contract quality thresholds with remedies, change control, voice rights, and transition assistance.
Frequently Asked Questions
What is the single most useful question to ask a localization vendor?
How they handle a re-cut of already-delivered content. A capable vendor diffs the source at segment level, regenerates only what changed, and preserves reviewed work everywhere else. A weak one re-does the asset and charges again. Since content revision is continuous in most organisations, the answer predicts the ongoing cost of the relationship more accurately than the rate card does.
Should we run a paid pilot?
Yes, and structure it into the RFP rather than adding it afterwards. Give every shortlisted vendor identical real content — not their choice of sample — and evaluate the outputs blind using your own native reviewers against a written scorecard. A pilot on your actual content is worth more than the entire written response, and paying for it gets you the vendor's normal process rather than a special effort.
Who should own the translation memory and glossaries?
You should, with a contractual right to export them in a documented, non-proprietary format within a defined period after termination and at no additional charge. These assets accumulate the value of every review cycle you have paid for. Leaving ownership unstated is the most common and most expensive omission in localization contracts.
How should we compare vendors quoting in different units?
Ask each for a worked end-to-end price on the same specific representative asset into the same three languages, with everything included. Per-source-minute, per-word, and per-asset pricing are not directly comparable, and the exercise also surfaces which costs sit outside each vendor's headline rate — setup, voice creation, extra review cycles, re-cut handling, and export fees.
Is it a problem if a vendor uses AI heavily?
No, and every vendor does now. What matters is whether they are straightforward about it and whether human review is placed where the risk actually is — regulated content, brand-critical material, languages where errors are hardest to detect. Evasiveness about the degree of automation is a more meaningful negative signal than the automation itself.
What should we do before writing the RFP at all?
Decide what quality means for your content, who will review it, and what your revision cycle looks like. A significant share of failed localization engagements fail on the buyer's side, because these were never settled and no vendor can settle them for you. Writing the scope section honestly tends to expose whichever of them is missing.
Related reading: Build vs Buy Video Localization | Localization Agency Video Services | Video Translation Pricing Models



