Understanding WCAG 1.2.2 Captions (Prerecorded)
What is it?
WCAG 1.2.2 Captions (Prerecorded) is a Level A criterion. Any prerecorded video that has audio needs captions. Captions carry everything the audio does: the dialogue, who’s speaking, and the sounds that matter. If someone can’t hear the audio, the captions should tell the same story.
Captions aren’t subtitles
Subtitles assume you can hear the audio and just need the words, often translated into another language. Captions assume you can’t hear it at all, so they include the non-speech sounds too. For accessibility, you want captions.
Subtitles: “Hello…” Captions: “[phone ringing] Hello… [tense music]”
Why it matters
Captions are how a lot of people follow video. Deaf and hard-of-hearing people rely on them. But they help far more people than that: anyone watching with the sound off, anyone in a loud or quiet space, and anyone who reads a language more easily than they hear it.
Who it affects
People who…
- Are d/Deaf or hard of hearing
- Watch with the sound off, on a train or in an open office
- Are in a noisy place where the audio is hard to make out
- Are more comfortable reading the language than hearing it
- Process spoken audio differently
Common failures
- No captions at all. The video goes up with audio and nothing else. Anyone who can’t hear it gets a silent, confusing experience with no way to follow along.
- Auto-captions left uncorrected. Auto-generated captions are a starting point, not a finish line. They mangle names, drop punctuation, miss technical terms, and skip who’s speaking.
- Skipping the sounds that matter. Captions that cover speech but ignore meaningful sounds leave gaps. A phone ringing, a door slamming, the music that sets the tension. If it carries meaning, it belongs in the captions.
- No speaker labels when it’s ambiguous. When someone off-camera speaks, or in hectic scenes with several people speaking at once, it’s unclear who said what.
- A transcript instead of captions. Transcripts are helpful in a different way, but they make following along while watching the video nearly impossible.
Solution
Caption every prerecorded video
Every prerecorded video with audio gets captions. That’s the baseline, and it’s the fix for a video without them. Everything after this is about getting them right: accurate words, the sounds that matter, clear speakers, and captions that live on the video itself.
Start from auto, then correct
Auto-captions are fine as a first draft. Fix the wrong words, add the punctuation, and check names and technical terms. Review them carefully before you publish.
Caption the sounds, not just the speech
Put meaningful non-speech sounds in brackets: [phone ringing], [applause], [tense music]. If a sound tells the viewer something, the caption should too.
Name the speaker when the video doesn’t
When you can’t tell who’s talking from the picture, put the speaker’s name in the caption. In a caption file, it’s a simple label before the line. For example:
00:00:04.000 –> 00:00:06.000
[phone ringing]
00:00:06.000 –> 00:00:08.500
Alex: I’ll get it.
Provide both captions and a transcript
Captions sync to the video for people watching it. A full transcript is a separate text version, and it serves a different purpose and audience. Ideally, provide both. The transcript falls under a different success criterion, 1.2.3 Audio Description or Media Alternative (Prerecorded), a Level A criterion.
Closed vs. open captions
Closed captions come from a separate file, and the viewer turns them on or off. Open captions are burned into the video and always show. Closed is usually the better default: viewers can turn them on, resize them, reposition them, and select a font and color. Additionally, they can be turned off for anyone who doesn’t need or want them.
(Example: a closed caption sits in a caption box the viewer can toggle with a CC button, like “Look at this cat! [meow]”. An open caption is burned into the picture and always shows, styled as part of the video.)
Match the audio, profanity included
Captions have to match what people actually hear.
- If a word is bleeped in the video, mask it in the captions too.
- If it isn’t bleeped, write it exactly as spoken.
Bleeping the captions when the audio wasn’t bleeped hands caption users a censored, second-class version. Clean up both or neither. Never one and not the other.
The one exception
WCAG 1.2.2 has a narrow exception. If the video is itself an alternative to text that’s already on the page, and it’s clearly labeled as such, then captions aren’t required. That’s rare. Almost every video you publish needs them.
Note that live video has a separate requirement for captions under WCAG 1.2.4 Captions (Live), Level AA.
How to test
- Turn the sound all the way off and watch your own video. Can you follow all of it?
- Check that the captions are accurate, not auto-generated guesswork, and that they match the audio, profanity and all.
- Confirm they’re synced to the video, that meaningful sounds are captioned, and that speakers are named whenever it isn’t clear who’s talking.
Closing
Captions aren’t just for people who can’t hear. You caption for one group and you help a dozen: the muted commute, the quiet office, the person learning the language. When the sound is off, the captions carry the whole story.
Learn more
Read the full breakdown: a20y.com/wcag-plain-english/1-2-2-captions-prerecorded
Want more clear and actionable WCAG breakdowns? Check out: wcagInPlainEnglish.com