Tutorials Accessibility

Beyond Auto-Captions: Caption Quality, Styling, and Readability

Beginner · ~25 min

Overview

Auto-captions solved the "no captions" problem and created a new one: captions that exist but fail the people who depend on them, misheard names, no speaker identification, three-line walls of text flashing by too fast to read. This is the quality-focused sequel to the site's guide on adding captions and subtitles: what separates captions that technically exist from captions that actually work.

What You Need

  • An auto-generated caption file as your starting draft
  • A caption/SRT editor (the site's SRT Editor works in the browser)
  • Fifteen minutes per content-hour budgeted for the correction pass

Steps

1

Understand why auto-captions alone don't cut it

Auto-captions reliably miss exactly what matters most: proper names, technical terms, numbers, who's speaking, and the non-speech sounds that carry meaning (a door slam, sarcastic laughter). The words they get right are the cheap ones. The errors cluster on the expensive ones.

2

Do the correction pass that matters

Triage, don't perfect: fix names, numbers, jargon, and any error that flips meaning (dropped "not"s are notorious). This pass takes minutes per content-hour and removes the errors that actually mislead, chasing every filler-word discrepancy is where caption budgets go to die.

3

Respect line length, duration, and reading speed

The fundamentals: roughly two lines maximum, around 32 to 42 characters per line, on screen long enough to actually read, viewers read slower than editors think. Break lines at natural phrase boundaries, not mid-clause, and never let a caption flash shorter than about a second.

4

Add speaker identification and sound cues

When speakers aren't visually obvious, label them ([MAYA] or a name-dash convention). Add the non-speech audio that carries meaning ([phone buzzes], [crowd groans]) and skip decorating with sounds that don't. This is the layer auto-captions almost never provide and deaf viewers most need.

5

Style for legibility

High contrast (light text, dark backing band or strong outline), a clean sans-serif at a size that survives a phone screen, positioned clear of on-screen text and graphics. Check placement against the site's Safe Zones Guide. Style consistency across your content matters more than any individual choice.

6

Choose open vs. closed captions deliberately

Closed captions (a toggleable track) are the accessibility default: viewers control them, players style them, and translations can be added. Open captions (burned into the picture) win for short-form social where platform players are unreliable. Many workflows ship both: closed on long-form, open on the vertical cuts.

Pro Tips

  • Build a correction glossary of your recurring names and jargon: most auto-caption errors repeat, and a find-and-replace list fixes them in bulk.
  • Watch a few minutes of your captioned video with the sound off. It's the fastest honest test of whether the captions alone carry the content.
  • Keep the corrected caption file as your master transcript. It feeds show notes, translations, and search, so the quality pass pays off repeatedly.

Captions Serve Far More People Than the Rules Require

Beyond deaf and hard-of-hearing viewers, captions serve commuters watching muted, non-native speakers, viewers in noisy rooms, and everyone skimming, on short-form platforms a large share of viewing happens sound-off. That's why caption quality is simultaneously an accessibility obligation and one of the highest-leverage retention improvements available: bad captions lose viewers who would have stayed.

Reading Speed Is the Invisible Constraint

Most caption failures aren't accuracy failures. They're timing failures. Text that's perfectly correct but on screen too briefly is functionally missing for the people who need it. Line length, duration, and phrase-boundary breaks are the craft that makes accurate text actually readable, and it's the part auto-generation handles worst.

Caption Formats, and Which One to Deliver

Caption files come in several formats and the differences matter mainly at delivery.

SRT is the simplest and most widely accepted: timings and text, essentially no styling. It is the safe default for uploading to platforms, and its limitations rarely matter because platforms apply their own styling anyway.

WebVTT is the web standard, supports positioning and basic styling, and is what you use for HTML5 video on your own site.

TTML and its broadcast profiles carry rich styling and positioning and are what broadcasters and some streaming platforms require in delivery specifications.

Burned-in captions are rendered permanently into the picture. They guarantee appearance and are common on social video where autoplay is muted, but they cannot be turned off, translated, resized by the viewer, or read by assistive technology, so they should supplement a real caption file rather than replace it.

What Automatic Captioning Still Gets Wrong

Speech recognition has improved enormously, which makes its remaining failure modes easier to miss because the output looks confident and mostly right.

Proper nouns and jargon are the most common errors, names, brands, technical terms, and acronyms are exactly the high-information words whose corruption most damages meaning.

Punctuation and sentence boundaries are inferred rather than heard, and a misplaced full stop changes meaning silently.

Speaker changes are frequently missed entirely, turning a conversation into an undifferentiated block of text where it is impossible to tell who said what.

Non-speech audio is simply absent, because the system transcribes speech and does not describe sound.

Accuracy also degrades with accent, overlapping speech, and background noise, meaning it performs worst on exactly the material where captions matter most.

Speaker Identification and Non-Speech Information

The distinction between a transcript and a caption file lives here. A caption serves someone who cannot hear the audio, so it has to carry the information the audio would have given them.

Speaker identification matters whenever more than one person speaks and the speaker is not obvious on screen. The convention is a name or role followed by a colon, used on speaker change rather than on every caption.

Non-speech information is written in square brackets and included when it carries meaning: [door slams], [laughter], [ominous music], [phone rings offscreen]. The test is whether a hearing viewer would take meaning from the sound. Ambient noise that carries nothing does not need describing, and over-annotating is its own readability problem.

Music deserves particular attention. If lyrics matter, caption them. If the music sets tone, describe the tone. "[music]" tells a viewer almost nothing.

A Caption QA Pass That Takes Ten Minutes

Most caption defects are catchable with a short structured check rather than a full re-read.

Play the video with sound off and captions on, at normal speed. This is the actual user experience and it exposes timing and reading-speed problems immediately, anything you cannot finish reading before it disappears is too fast regardless of what the numbers say.

Then check specifically: do captions appear before the speech they transcribe, or lag behind it. Do they cover on-screen text, faces, or lower thirds. Do speaker labels appear at every change. Are proper nouns spelled correctly. Do line breaks fall at sensible grammatical points rather than mid-phrase. And does any caption run to three or more lines.

Line breaking is the most-neglected item and one of the most damaging. Breaking a line mid-phrase forces the reader to hold an incomplete thought while the eye travels, which measurably slows reading and is why professionally captioned material breaks on clause boundaries.

Translation, Multi-Language Subtitles, and What Not to Automate

Once captions exist, translation looks like a cheap next step, and it is, provided you understand which parts genuinely automate.

Machine translation of an accurate caption file produces usable results for straightforward informational content, and is dramatically better than no subtitle at all for reach. It degrades predictably in three places: idiom, which translates literally and lands as nonsense. Technical terminology, where the domain-correct term differs from the general one. And humour, which frequently depends on structure that does not survive.

The quality of the source file governs everything downstream. Translating an uncorrected automatic transcript compounds every error, a misheard proper noun becomes a confidently mistranslated one, which is why the correction pass on the original language is the highest-leverage work in a multi-language pipeline.

Reading speed also changes between languages. Text that expands significantly in translation will not fit the timings that worked for the original, so translated files frequently need timing adjustment rather than a straight text swap.

For anything where precision matters (legal, medical, safety, or brand-sensitive material) machine translation should be a first draft reviewed by someone who speaks the language. The failure mode is not gibberish, which is obvious, but fluent text that says something subtly different from the original.

FAQ

Q: How accurate is accurate enough for captions?
A: The practical bar: every name, number, and technical term correct, and no error that changes meaning. A stray filler-word discrepancy doesn't hurt anyone. A misheard name or a dropped "not" does. Auto-captions typically land most of the words but miss exactly the high-stakes ones: which is why the correction pass focuses on names, numbers, jargon, and negations first.

Q: Should captions be verbatim or cleaned up?
A: For accessibility, lean verbatim: deaf and hard-of-hearing viewers deserve the same content as everyone else, including tone-carrying repetitions and false starts where they matter. Light cleanup of pure stutters and fillers is widely accepted. Rewriting or paraphrasing what was said is not. If in doubt, keep what carries meaning.

Q: How many characters per line should a caption have?
A: Around 32 to 42 characters per line, with a maximum of two lines on screen at once, is the widely used range. The constraint is reading speed rather than the character count itself. The numbers exist because they keep a two-line caption readable within the time it is typically displayed.

Q: Should I burn captions into the video or use a caption file?
A: Use a caption file wherever the platform supports one, and add burned-in captions as a supplement for muted-autoplay social video. Burned-in text cannot be disabled, translated, resized, or read by assistive technology, so it should never be the only provision on a platform that accepts a real file.

Q: Do I need to caption music and sound effects?
A: Caption them when they carry meaning. Lyrics that matter should be transcribed. Music that establishes tone should be described in terms of that tone rather than labelled generically. Sound effects need describing when a hearing viewer would draw information from them. An offscreen crash, a phone ringing, a door closing.

Q: Can I just machine-translate my captions into other languages?
A: As a first pass, yes, and it substantially increases reach for straightforward informational content. Correct the source-language file first, because errors compound through translation. Have a speaker review anything where precision matters, since the risk is not obvious gibberish but fluent text that means something subtly different. Expect to adjust timings, as text length changes between languages.

Translate this page

Machine translation provided by Google Translate, on Google’s servers. We do not check these translations and they will get technical terms wrong. The English page is the authoritative one. Following a link sends this page’s address to Google. Your browser may also offer to translate this page itself, which keeps the request on your device.