Person reading at a desk with a notebook

The Audience Nobody Designs For

Accessibility conversations about video default to sensory access: can the viewer hear it, can the viewer see it. Those are the requirements with clear criteria and clear deliverables.

Cognitive accessibility is the part that has no checkbox. It covers everyone for whom the content is technically perceivable but practically hard to follow: people with learning disabilities, people with attention or memory conditions, people under stress or fatigue, older viewers, and — very large in any global content operation — people watching in a language they speak as a second or third language.

That last group is why this belongs in a localization conversation rather than only an accessibility one. A significant share of the audience for any multilingual video programme is watching in a language they did not grow up with. Content that assumes native-speaker fluency in idiom, pace, and structure loses them.

The useful thing about cognitive accessibility is that the work overlaps almost perfectly with what makes content translate well. Simplifying structure for comprehension also removes exactly the constructions that produce translation errors. Two problems, one intervention.

What Makes Video Hard to Follow

Video has properties that print does not, and they cut against comprehension.

It moves at the presenter's pace, not the viewer's. A reader who loses the thread rereads a sentence without thinking about it. A viewer must find the control, scrub back, and re-enter the flow — a deliberate act most people do not perform. Anything missed is usually missed permanently.

It demands parallel processing. Audio, on-screen text, and visual demonstration arrive simultaneously. Where they carry different information, the viewer must divide attention. Where captions are on, they must also read.

It offers no structural overview. A document shows headings, length, and shape at a glance. A video reveals its structure only by being watched.

It has no built-in reference. A reader flips back to a definition. A viewer cannot, so terms introduced once must be remembered.

These properties are why the same content in video form is often harder to follow than in written form, and why the fixes are mostly structural rather than cosmetic.

Script-Level Changes That Work

Most cognitive accessibility for video is decided in the script, before anything is filmed.

Say what the video covers at the start. Thirty seconds of orientation — what this is, who it is for, what you will know at the end — gives the viewer a frame to hang the rest on. It also lets people who need something else leave early, which is a service to them.

One idea per sentence. Compound sentences with subordinate clauses are where both comprehension and translation break down. Splitting them costs a few words and removes the ambiguity about which clause attaches to what.

Prefer active voice and explicit subjects. "The system sends a confirmation" is easier than "a confirmation is sent" and it survives translation into languages with different information structure far better, because the translator does not have to reconstruct who did what.

Introduce a term before using it. Define once, clearly, then use consistently. Do not use two words for one concept for the sake of variety — this is good writing advice for essays and bad advice for instructional content in any language.

Avoid idiom, metaphor, and cultural reference. "Circle back", "low-hanging fruit", "out of left field" — these are opaque to second-language viewers, unpredictable for translators, and add nothing.

Signpost transitions. "That covers setup. Next, permissions." Explicit transitions replace the structural cues that video lacks.

Number things. "There are three steps" followed by "first", "second", "third" gives the viewer a progress model and a memory aid.

Keep numbers and figures simple, and repeat them on screen. Spoken figures are hard to retain. If a number matters, show it.

Notebook, pen, and laptop on a desk

Pacing and Structure

Slow down slightly at the load points. Delivery does not need to be uniformly slow — that is tedious and does not help. It needs to slow at the moments where new information arrives, and it can move briskly through recap and transition.

Leave pauses after complex points. A beat of silence is processing time. It is also where audio description goes if it is ever needed, so the pauses do double duty.

Cut content into short segments with clean boundaries. A forty-minute continuous video is a worse learning object than eight five-minute ones covering the same ground, because it offers no re-entry points. Segmentation also makes translation review tractable, since a reviewer can work through discrete units.

Recap before moving on. A one-sentence summary at the end of a segment consolidates what the viewer just heard and gives anyone who drifted a chance to rejoin.

Match the visual to the audio. Where the screen shows one thing and the narration discusses another, the viewer must choose. Where they agree, the visual reinforces rather than competes.

Keep on-screen text short and let it stay long enough to read. Text that appears and disappears in two seconds serves nobody, and it becomes actively hostile once translated, since target-language text is usually longer.

Why This Makes Translation Cheaper

The overlap between plain language and translatable language is not partial. It is nearly complete.

Short sentences reduce ambiguity. Machine translation errors cluster in long sentences with multiple clauses, because the system must decide what attaches to what and there is often no unambiguous answer. Shorter sentences remove the decision.

Explicit subjects survive language change. Languages differ in how much they allow ellipsis. A source sentence that omits the subject forces the translator or the system to guess, and languages requiring an explicit subject will get it wrong some of the time.

Consistent terminology enables glossaries. Terminology extraction and locking work only when the source uses one term per concept. Synonym variation in the source multiplies into inconsistency across every target language.

No idiom means no idiom failures. Idiomatic expressions either translate literally into nonsense or require a creative equivalent that a reviewer must supply. Removing them removes an entire error class.

Simple structures reduce text expansion pressure. Verbose source text expands further in translation, and the languages that expand most are exactly the ones where subtitle reading speed is already tight. Concise source gives the target room.

Deliberate pacing gives dubbing headroom. Dubbed audio in an expanding language needs time. A source track delivered at a relentless pace leaves no slack, forcing either compressed delivery or drift against the picture. Natural pauses in the source become the room the target language uses.

The practical implication is that a plain-language editing pass on the source script pays for itself twice: once in comprehension for the source audience, and again in every target language, where it reduces review effort and error rates simultaneously.

Testing Comprehension

Readability formulas designed for print — grade-level scores and the like — are weak instruments for spoken content and worse for translated content. They measure sentence and word length, which correlate with difficulty but do not capture structure, pacing, or terminology load.

More useful signals:

Ask someone outside the domain to summarise it. If they cannot state the main point after one viewing, the structure is not working.

Test with second-language viewers. They surface idiom, pace, and assumed knowledge faster than anyone else, and they are a large share of the real audience.

Watch where people scrub back. Player analytics showing repeated rewinds at a specific timestamp are a comprehension failure marker with a timecode attached.

Track drop-off. A sharp fall at a particular point usually means something got hard, not that the topic got boring.

Read the caption track alone. If the words do not make sense without the picture, the audio is doing more work than it should, and that gap will widen in translation.

Group watching a presentation on a screen

A Working Checklist

  • Open with thirty seconds stating what the video covers and who it is for.
  • Write one idea per sentence and split compound sentences.
  • Use active voice with explicit subjects.
  • Define each term once and then use it consistently, without synonym variation.
  • Remove idiom, metaphor, and culture-specific reference from instructional content.
  • Signpost every transition explicitly.
  • Number sequences and state the count up front.
  • Show important numbers on screen as well as saying them.
  • Slow delivery at points of new information and leave pauses afterwards.
  • Segment long content into short units with clean boundaries.
  • Recap each segment in one sentence before moving on.
  • Keep the visual and the audio aligned rather than competing.
  • Leave on-screen text up long enough to read at target-language length.
  • Run a plain-language editing pass on the source script before translation, not after.
  • Test comprehension with second-language viewers and watch scrub-back analytics.

Frequently Asked Questions

Is cognitive accessibility a WCAG requirement?

Only partially. WCAG includes some criteria that touch it — reading level appears at Level AAA, and there are requirements around consistent navigation and predictable behaviour — but there is no Level AA criterion that mandates plain language in video narration. Treat it as a quality and reach issue rather than a compliance one. The absence of a checkbox is precisely why it gets skipped.

Does simplifying language make content sound patronising?

Done badly, yes. The failure mode is removing precision along with complexity, which produces vague content that says less. Done well, plain language keeps every technical term the content actually needs, defines them, and simplifies the sentences around them. Specialists generally prefer clear technical writing to dense technical writing; what reads as patronising is usually over-explanation, not short sentences.

How much does a plain-language pass slow down production?

An hour or two per finished video minute for an editing pass on the script, done before filming. It is recovered several times over in translation review, because the error classes it eliminates — ambiguous attachment, inconsistent terminology, idiom — are exactly the ones that need human correction in every target language. For a programme running eight languages, the arithmetic is not close.

Does this apply to entertainment content or only instructional content?

Mostly instructional, informational, and commercial content, where comprehension is the point. Drama and entertainment have different goals, and flattening a script's voice to improve translatability would be the wrong trade. Even there, though, the structural advice about on-screen text duration and pacing headroom for dubbing still holds.

Should we simplify the translated version or the source?

The source. Simplification applied only to targets means every language pays the cost separately and the results diverge. Fixing the source once propagates the benefit to every language, including the original.

How does this interact with subtitle segmentation?

Directly, and favourably. Subtitle cues should break at syntactic boundaries, and a script written in short single-clause sentences supplies those boundaries naturally. Long compound sentences force segmentation tools to break mid-clause, producing cues that read awkwardly and that get worse in translation as the target-language string lengthens. The same editing habit that improves comprehension therefore also improves the caption track, in every language, without any additional work.

Does on-screen text need the same treatment?

More so, because it competes with the narration for attention rather than reinforcing it. On-screen text should be short enough to read at target-language length, use the same terminology as the audio rather than a paraphrase, and stay on screen long enough for a second-language reader. Text that says something different from what the narrator is saying at that moment is the worst case: the viewer must choose which to process, and most process neither well.

What is the single highest-value change?

Splitting long sentences. It improves comprehension for second-language and cognitively diverse viewers, removes the largest single source of machine translation error, reduces text expansion, and makes subtitle segmentation cleaner — all from one editing habit.


Related reading: Video Accessibility Guide | Text Expansion in Translation | Video Translation Common Mistakes