Skip to main content

Glossary

Transcription

The automatic conversion of a recording's speech into text, which becomes the basis for captions, search and summaries.

Transcription turns the speech in a recording into text. In video tools it is almost always automatic, produced by a speech recognition model within a minute or two of the recording finishing, and it is the foundation that several other features are built on.

Its most visible use is captions, but the transcript is more valuable than that suggests. It makes a video searchable, which is the difference between a library of recordings and an archive nobody can use. It gives accessibility tooling something to work with. And it is the input for summaries, chapters and translation, none of which can exist without it.

Accuracy depends far more on the audio than the model. Clear speech in a quiet room transcribes almost perfectly; a echoing room, crosstalk or heavy background noise degrades it quickly. Specialist vocabulary and product names are the other common failure, which is why reviewing the transcript of anything public is worth the few minutes it takes.

For work video the transcript is often the artefact colleagues actually consume. Many people would rather skim two hundred words than watch four minutes, and offering both is what makes a recording usable by a team that spans time zones.

Editing the transcript is worth the few minutes when the recording matters. Correcting product names and specialist terms fixes the captions, the summary and any translation at the same time, because all three are generated from it. That single pass is the highest-value review anybody can do on a recording they intend to publish.

In Zidi

Zidi transcribes every recording automatically, and that transcript is what the captions, the summary, the chapters and the translations are all generated from. It is available to read alongside the video rather than only inside the player.

Back to the full glossary.