Overview
Video accessibility crossed a line in June 2025: with the EU's Accessibility Act in force, accessible publishing became a legal baseline for many services rather than a best practice, and other jurisdictions have their own rules moving the same direction. This guide is the practical checklist: what the accessibility stack actually consists of, how to audit what you already publish, and how to build it into the workflow so it stops being a retrofit. Scope and specifics vary by jurisdiction and service type. Verify what applies to you. This is not legal advice.
What You Need
- An inventory of where your video is published (platforms, your own site, embeds)
- Your caption/transcript workflow, honestly assessed
- A view of which jurisdictions and rules apply to your audience and service
Steps
Know the accessibility stack
Four layers, each serving different users: captions (deaf/hard-of-hearing, sound-off viewers), audio description (blind/low-vision viewers), a published transcript (screen-reader users, skimmers, search), and an accessible player (keyboard and screen-reader operable). "Accessible video" means the stack, not just captions.
Understand the compliance-era baseline in general terms
The general shape across jurisdictions: in-scope services must provide captions and audio description, in the language of each market served, with the EU's rules in force since June 2025 and existing libraries facing deadlines later in the decade. Which rules bind you depends on what you operate and where your audience is. This is exactly the kind of detail to verify rather than assume.
Audit what you already publish
Spot-check honestly: pick recent and top-traffic videos and check each stack layer, are captions corrected or raw auto-captions, does AD exist at all, is there a transcript, does the player work by keyboard? Most audits find captions-only, which makes the gap list short and concrete.
Prioritize a back catalogue honestly
Nobody remediates 500 videos at once. Triage by traffic and lifespan: top-viewed and evergreen content first, time-limited and low-traffic content last or never (documented as such). A written, dated plan beats both denial and heroics.
Check your players and embeds
Major platform players handle the basics (keyboard control, caption rendering, screen-reader labels) reasonably well, self-hosted and custom players are where accessibility usually breaks. If you embed or self-host, test: can you operate every control by keyboard alone, do captions render, does a screen reader announce the controls sensibly?
Build accessibility into the workflow, not as a retrofit
The sustainable version: the corrected transcript from your edit (see caption quality) becomes captions and the published transcript. AD gets written against the locked cut. The publish checklist includes the stack. Done in-line, it's a modest per-video cost. Done as an annual cleanup, it's a project everyone dreads.
Pro Tips
- Add the accessibility stack to your publish checklist template once, per-video decisions are where consistency dies.
- Keep dated records of what was remediated when: if compliance is ever questioned, a documented plan and progress trail is your best position.
- Watch your own top video with captions on, sound off, then listen to it with the screen off. The two crudest tests catch the most real problems.
Knowledge Base
Why the Baseline Moved
Accessibility rules for video used to concentrate on broadcasters. The current generation of rules, the EU's Accessibility Act most prominently, extends expectations to streaming services, e-learning, and video published as part of consumer-facing digital services. The practical consequence for creators and small teams: accessibility questions that used to be someone else's department are now workflow questions, which is exactly why building the stack into the routine beats treating it as a compliance scramble.
Accessibility and Reach Are the Same Investment
Every layer of the stack serves a mainstream audience too: captions serve sound-off viewers, transcripts serve search and skimmers, legible design serves everyone on a phone in sunlight. The compliance framing gets budgets approved. The reach framing is why the work pays for itself even where no rule applies.
The Conformance Levels, and Which One You Are Actually Aiming At
Accessibility guidance is organised into three conformance levels, and knowing which one applies saves a great deal of argument. Level A is the minimum, without it, some people simply cannot use the content at all. Level AA is the level nearly every regulation, procurement requirement, and institutional policy actually references, and it is the sensible target for published video. Level AAA is aspirational and not expected across a whole site.
For video specifically, the practical AA requirements are captions for prerecorded audio, audio description for visual information not conveyed in the soundtrack, controls that work by keyboard, sufficient contrast on any text or interface you overlay, and no content that flashes in a way likely to trigger seizures.
The useful reframing is that AA is not a ceiling you strain toward but a floor that most published video currently sits below. Captions alone, done properly, move a large share of content most of the way there.
Captions, Subtitles, and Audio Description Are Three Different Things
These get used interchangeably and cover genuinely different needs, which is why "we added subtitles" is not always an answer to an accessibility requirement.
Captions assume the viewer cannot hear the audio. They therefore include speaker identification and meaningful non-speech sound (a door slamming, music becoming ominous, a phone ringing offscreen) because that information exists only in the soundtrack.
Subtitles assume the viewer can hear but does not understand the language. They translate dialogue and generally omit non-speech information, because the viewer can hear it themselves.
Audio description assumes the viewer cannot see the picture. A narrator describes visual information the soundtrack does not carry, who entered the room, what the chart shows, what a character is doing during a silence. This is the requirement most often skipped, and the one that matters most for instructional and data-heavy video, where meaning frequently lives on screen and nowhere else.
Build It Into the Workflow Instead of Retrofitting
Accessibility is cheap when it is part of production and expensive when it is a task added after delivery. The difference is largely about when information is still available.
A script that exists before the shoot is most of a caption file already. A presenter who has been asked to say what is on screen rather than "as you can see here" removes much of the need for audio description. Designing lower thirds and on-screen text with adequate contrast at the template stage costs nothing. Fixing contrast across an existing library costs a redesign.
The retrofit path is worse in every dimension: someone transcribes speech that was already written down, describes visuals whose intent nobody recorded, and re-exports graphics to fix contrast decisions made months earlier. Teams that treat accessibility as a delivery checkbox keep paying this cost on every project, which is why it gets perceived as expensive.
Where the Legal Requirements Come From
This is not legal advice, but the shape of the landscape is worth knowing because it determines who is obliged rather than merely encouraged.
Obligations generally attach to who is publishing and to whom, not to the video itself. Public sector bodies are the most consistently covered, typically through regulations that reference the AA level directly. Education, broadcast, and companies selling into regulated markets are commonly covered through procurement requirements even where no statute names them.
The trend across jurisdictions has been toward broader coverage of commercial services rather than narrower, and toward accessibility being specified in contracts as a deliverable. The practical consequence for a creator or small studio is that the question is increasingly asked by clients rather than regulators, and being able to answer it is a commercial advantage before it is a compliance matter.
Test With Real Assistive Technology, Not Just a Checklist
Automated checkers and conformance checklists catch the objective failures (missing captions, insufficient contrast, absent labels) and they cannot tell you whether the result is actually usable. That gap is where most technically-compliant-but-unusable video lives.
The cheapest meaningful test is to use your own content the way someone else would. Watch a video with the sound off and read only the captions: does it still make sense, or do the captions describe speech while the meaning lives in visuals nobody described? Then listen without watching: does the audio alone convey the content, or is it a series of references to things on screen?
Next, operate your player using only the keyboard. Tab to the controls, play, pause, adjust volume, enable captions, and go fullscreen without touching a pointer. Custom players frequently fail here while looking fine, because the controls were built as clickable graphics rather than as buttons.
If you can, try a screen reader. Every major operating system includes one at no cost. The experience is unfamiliar and initially awkward, and half an hour with it will teach you more about your own pages than any checklist.
The recurring lesson is that compliance and usability are related but not identical, and only the second one actually serves anybody.
FAQ
Q: I'm a solo creator outside the EU (does any of this apply to me?
A: Legally, maybe) accessibility rules exist in many jurisdictions and some apply based on where your audience is, not where you are. Verify your own situation. Practically, definitely: captions, transcripts, and legible design expand your actual audience and improve retention regardless of any mandate, and platform norms are moving the same direction the laws are.
Q: Where do I start with a 500-video backlog?
A: Traffic-weighted triage: your top-viewed and evergreen videos serve the most people per hour of remediation work, so caption and transcript those first. New content gets the full treatment from day one so the backlog stops growing. A dated, prioritized plan is also far more defensible than either ignoring the catalogue or pretending you'll fix everything at once.
Q: Are auto-generated captions enough to be compliant?
A: Generally no. Automatic captions typically miss punctuation, speaker changes, and non-speech audio, and their accuracy falls sharply with accents, overlapping speech, technical vocabulary, and background noise. They are a useful starting draft that a human corrects, which is far faster than transcribing from scratch, but shipped uncorrected they usually fail the standard they are meant to meet.
Q: Do I need audio description for every video?
A: You need it where visual information is not otherwise conveyed. A talking-head video where the speaker says everything that matters may need none. A tutorial that demonstrates on screen while saying "click here" needs it badly. The cheapest fix is often to change the narration so the visual information is spoken naturally, which removes the need for a separate description track.
Q: What contrast ratio do I need for on-screen text?
A: At AA, normal text needs a contrast ratio of at least 4.5:1 against its background, and large text at least 3:1. Text over video is the difficult case because the background moves. The reliable solution is a solid or heavily-shaded plate behind the text rather than relying on the footage staying dark enough.
Q: Is passing an automated accessibility checker enough?
A: No. Automated tools catch objective failures such as missing captions, poor contrast, and unlabelled controls, but they cannot judge whether captions are accurate, whether visual information is conveyed some other way, or whether the player is genuinely operable. Watch your content muted, listen to it unwatched, and operate the player by keyboard. Those three checks find what tools miss.
Translate this page
- Español
- 简体中文
- हिन्दी
- العربية
- Português
- Français
- Deutsch
- 日本語
- Русский
- Bahasa Indonesia
- 한국어
- Italiano
- Türkçe
- Tiếng Việt
- Polski
- Nederlands
Machine translation provided by Google Translate, on Google’s servers. We do not check these translations and they will get technical terms wrong. The English page is the authoritative one. Following a link sends this page’s address to Google. Your browser may also offer to translate this page itself, which keeps the request on your device.