Glossary
Captions versus subtitles
Captions render speech and relevant sound for viewers who cannot hear it; subtitles translate dialogue for viewers who cannot follow the language.
The two words are used interchangeably in marketing and mean different things in practice. Captions are written for viewers who cannot hear the audio and therefore include relevant non-speech sound. Subtitles assume the viewer can hear but does not follow the language, so they translate dialogue and leave sound cues out.
There is a second distinction between open and closed. Closed captions are a separate track the viewer can switch on and off, which is the better default because it leaves the choice with them and keeps the text searchable. Open captions are burned into the picture and cannot be turned off, which is what social platforms effectively require since most feeds play silently.
For work video, captions are not an accessibility afterthought. A large share of video in an office is watched with the sound off, and a recording with no captions is unwatchable in an open-plan room, on a train, or in a meeting the viewer is half-attending.
Both are generated from the transcript, which is why caption quality is really transcription quality, and why reviewing the transcript once fixes both at the same time.
Placement and timing decide whether captions help or annoy. Text that lags the speech by more than a moment is harder to follow than no text at all, and captions that sit over the part of the frame carrying the information force the viewer to choose between reading and watching. Most players allow a position adjustment, and it is worth making once for recordings that will be watched widely.
There is a legal dimension in some contexts, which is worth knowing even though it is rarely the reason teams add captions. Accessibility requirements in several jurisdictions cover video published by public bodies and, in some readings, by large employers. Adding captions because it makes recordings watchable in an open-plan office happens to satisfy most of what those rules ask for.
In Zidi
Zidi generates captions from the transcript and can translate them, so one recording can carry text in more than one language.
Related terms
Further reading
Back to the full glossary.