Accessibility

Captions are not decoration. They are part of the content.

A transcript, a subtitle track, and an accessibility caption file may share words, but they solve different problems. Publishing the right one requires timing, sound information, readable layout, and review.

The short version

  • Captions reproduce speech and meaningful non-speech audio in sync with video; translated subtitles serve a different primary purpose.
  • WCAG requires captions for prerecorded synchronized media at Level A and live synchronized media at Level AA, subject to the standard’s scope and exceptions.
  • Automatic output is a draft when names, timing, speaker changes or important sounds matter.
  • Good accessibility also means readable lines, safe positioning, keyboard operation and a transcript that can be searched and navigated.

1. Captions, subtitles, and transcripts are related—not interchangeable

Captions are synchronized text for the speech and non-speech audio needed to understand a video. They may identify the speaker and describe a meaningful sound. Subtitles often refer to a translation for viewers who can hear the original audio. A transcript presents the content as a document and can support audio-only material, reading, search, quotation, and navigation.13

One timed transcript can help produce all three, but the final forms need different editing. A transcript can carry paragraphs and fuller context. A caption needs short, readable cues that appear at the right moment and do not cover essential visual information.

2. What the web accessibility standard actually says

W3C’s guidance explains that prerecorded video with necessary audio needs captions at WCAG Level A, while live synchronized media needs captions at Level AA. Captions include dialogue, speaker identification, and meaningful sound effects—not speech alone. The exact legal obligation depends on jurisdiction and context; WCAG guidance is a technical standard, not individualized legal advice.21

The distinction matters because a silent transcript link beside a video is not the same experience as synchronized captions. Conversely, captions inside a player may not give an audio-only user the structured document they need. Accessible publishing starts by identifying the media and the user need before choosing the format.

3. The audience is larger than a compliance checklist

The World Health Organization reports that about 1.5 billion people live with some degree of hearing loss, including roughly 430 million who require rehabilitation services. Captions are also used in noisy rooms, quiet public spaces, classrooms, second-language viewing, and any situation where listening is difficult or inappropriate.41

Text creates additional access paths: it can be searched, quoted, translated, enlarged, read at a chosen pace, or used to jump to a moment in the recording. W3C notes that interactive transcripts can highlight the current phrase and let a reader select text to move playback to that point.1

4. Why automatic captions still need an editorial pass

Automatic speech recognition can remove most of the mechanical work, but it does not know which error is consequential. A wrong surname, speaker change, negation, or number can alter meaning. Important sound cues may not be present in the speech transcript at all. W3C’s practical guidance explicitly warns that automatically generated captions often need editing.1

Review against the audio, not only against a spellchecker. Check the beginning and end for missing content; verify names and numbers; mark speaker changes; add meaningful sound information; and watch the final video to confirm timing, line length, contrast, and safe placement.

5. A dependable publishing checklist

Keep a corrected master transcript with word- or segment-level timing. Derive caption and subtitle tracks from that master. Preserve the original media and revision history, and test the exported file in the player where people will actually watch it. A valid SRT file is not automatically a readable one.

KolWrite’s subtitle and transcript workflow is designed around that shared timed source: review the text beside playback, correct speaker and language details, then export the format the audience needs. Accessibility is achieved by the finished experience, not by the presence of an export button.

Common questions

Are subtitles and captions the same?

They may use the same file formats, but captions are intended to provide the speech and meaningful sounds for people who cannot hear the audio, while subtitles commonly translate dialogue for people who can hear it. Terminology varies by region.

Are automatic captions WCAG compliant without review?

Compliance depends on the final accuracy and completeness, not how the text was created. Automatic output can be a strong starting point, but names, timing, speaker identity, omissions and meaningful sounds should be reviewed.

Does a transcript replace captions for video?

Generally no. A transcript is valuable, but prerecorded synchronized video with necessary audio calls for synchronized captions under WCAG. Different media and exceptions should be evaluated against the standard and applicable law.

Sources and further reading

Primary research and standards used for this article. Links open the original publication.

  1. 1 W3C Web Accessibility Initiative — Captions/Subtitles
  2. 2 W3C — Understanding WCAG 2.2, Captions (Prerecorded)
  3. 3 W3C Web Accessibility Initiative — Transcripts
  4. 4 World Health Organization — Deafness and hearing loss
All research

Continue reading

Evaluation

Transcription accuracy is not one number

Word error rate is useful—but a transcript can score well and still fail the exact task you hired it to do. A fair comparison measures the recording, the workflow, and the cost of correction.

4 min read
Read article