File formats and audio

Captioning

Showing speech as readable text on screen, timed and limited in lines and characters.

·Also called: subtitling, captions

Captioning is showing speech as text on screen, synchronised with the audio.

It gets confused with transcription, but it is a different job. Transcription produces the text. Captioning makes it readable at the pace a video runs.

The rules that make text readable

  • A maximum of two lines at a time.
  • Around 37 to 42 characters per line. The eye cannot take in longer lines.
  • At least one second on screen, and rarely more than six.
  • Reading speed around 15 to 20 characters per second for adults. If the speech is faster, the text has to be shortened.
  • Break at natural boundaries, not in the middle of a phrase.

This is why captions are often not word for word: if the text is to be readable, it sometimes has to be tightened.

Subtitles or captions

In English the distinction is fairly clear:

  • Subtitles assume the viewer can hear the audio, and are often a translation into another language.
  • Captions also describe relevant sounds for viewers who cannot hear the audio, and are usually in the same language as the speech.

Many other languages use one word for both, and the distinction has to be made from context.

The file formats

SRT for general use, VTT for the web and wherever you need speaker names or positioning.

See also

SRT, VTT, real-time transcription.

Speech, written out

Sayable transcribes speech with dialect, tells the speakers apart and drafts the summary. Free to get started.