Glossary

Speech-to-textSTT

Speech-to-text (STT) is the technology that converts spoken audio into written words. It is the umbrella term covering live dictation, file transcription, captioning, and voice commands.

The conversion is performed by a speech recognition model, which may run in the cloud (audio is uploaded to a server) or on-device (audio never leaves the machine). The choice determines both privacy properties and offline behavior.

Raw STT output is a transcript: lowercase-ish, lightly punctuated, with fillers and false starts intact. Products built on STT differ mostly in what they do after recognition, from nothing at all to a full formatting pass.

Your voice was always faster.

Not out yet. $29 USD once when it is, with no account, no subscription and nothing uploaded.

No spam, no sharing. One email when Flit ships, and you can leave in one click.

macOS 14 or later · Apple silicon