Glossary

The terms, briefly explained

The field is full of abbreviations and words that mean something else in everyday speech. Here are 30 of them, explained plainly, short enough to finish.

Speech recognition and models

ASR
Automatic speech recognition, the technical term for turning speech into text automatically.
Code-switching
Shifting between two or more languages in the same conversation, often within a single sentence.
Diarisation
Splitting a recording into segments by who is speaking, without necessarily knowing who those people are.
Fine-tuning
Training a finished model further on a narrower material, for instance speech in one language.
Hallucination
Text the model produces without there being any basis for it in the audio.
Language model
A model that predicts likely word sequences, used both in transcription and to write summaries.
NB-Whisper
The National Library of Norway's further-trained version of Whisper, adapted to Norwegian speech and dialect.
Speaker recognition
Tying a voice to a known person, not merely telling it apart from other voices.
Speech recognition
The technology that turns the sound of speech into written text.
Voice activity detection (VAD)
The technique that decides where in an audio recording somebody is speaking, and where it is silent.
Whisper
An open speech recognition model from OpenAI that can be run on your own machine.
Word error rate (WER)
The share of words that are wrong, missing or added, divided by the number of words in the reference.

Methods and workflow

Batch transcription
Transcribing a finished recording afterwards, rather than while the speaking happens.
Cloud transcription
Transcription where the audio is uploaded and processed on the vendor's servers.
Local transcription
Transcription that runs on your own machine, without the audio being sent to a server.
Post-editing
Going through and correcting a machine-generated transcript against the audio.
Real-time transcription
Transcription that happens while the speaking happens, with one to three seconds of delay.

File formats and audio

Audio format
The file format the audio is stored in, for instance m4a, mp3 or wav.
Captioning
Showing speech as readable text on screen, timed and limited in lines and characters.
Sample rate
How many times per second the audio signal is measured, given in hertz.
SRT
The most common subtitle format: numbered text blocks with a start and end time.
Timestamp
A time reference tying a piece of the text to a particular point in the audio recording.
VTT
The web standard for subtitles, with support for styling, positioning and metadata.

Transcript forms

Edited transcript
A transcript where the content is kept, but phrased in whole, readable sentences.
Verbatim
Word-for-word transcription that includes hesitation, repetition, interruptions and often pauses.

Meetings and minutes

Action item
A concrete task following from a meeting, with one owner and one deadline.
Any other business
The standing last item on the agenda, for short matters not submitted in advance.
Decision log
A running list of decisions across meetings, with date, owner and reasoning.
Dissent
A member voting against, or recording disagreement with, a decision.
Statement for the record
A declaration a member requires to be entered in the minutes, regardless of what the majority thinks.

Why English dominates the field

Research on speech recognition was done in English, and the terms came along with it. Some languages have good native words, such as the Norwegian ordfeilrate for word error rate. Others do not, and then the English term beats inventing something nobody recognises. Where both are in use, both are listed here.

Speech, written out

Sayable transcribes speech with dialect, tells the speakers apart and drafts the summary. Free to get started.