Speech recognition went from mediocre to genuinely good in about five years, and most people never noticed the transition because it happened inside products they already used. Understanding how it works explains a lot about why it fails in the specific ways it does.
Step one: sound becomes numbers #
A microphone measures air pressure many thousands of times per second. Standard speech audio is sampled at 16,000 times a second, so one second of speech is a list of 16,000 numbers, each one the air pressure at that instant.
That list is the raw material. It is also nearly useless in that form, because the information a human hears is not in the individual pressure values, it is in the frequencies present over time.
Step two: the spectrogram #
So the first real processing step converts the waveform into a spectrogram, a picture of which frequencies are present at each moment.
The method is to take a short window of audio, twenty five milliseconds or so, run a Fourier transform on it to decompose it into frequency components, then slide the window forward by ten milliseconds and repeat. Stack the results and you get a two-dimensional image: time along one axis, frequency along the other, and brightness showing how much energy is at that frequency at that moment.
Then it gets adjusted to a mel scale, which compresses the frequency axis to match human hearing. We discriminate finely between low frequencies and coarsely between high ones, and the mel scale bends the axis accordingly so the representation spends its resolution where the information is.
The result is a mel spectrogram, and it is what nearly every modern speech system actually consumes. Speech recognition, at this point, has become an image problem.
This is worth sitting with, because it explains a lot of the field’s history. Vowels show up as distinctive horizontal bands. Consonants show up as bursts and noise. Different speakers produce recognizably different but structurally similar patterns for the same word.
Step three: the model #
Older systems did this in stages. An acoustic model mapped audio to phonemes, a pronunciation dictionary mapped phonemes to words, and a language model picked the most likely word sequence. Each piece was trained separately, each had its own failure modes, and errors compounded through the chain.
Modern systems are end to end. One neural network takes the spectrogram and produces text directly. No separate phoneme stage, no pronunciation dictionary, no hand-built pipeline.
The dominant architecture is a transformer encoder-decoder, the same family behind large language models. The encoder reads the spectrogram and builds a representation of it. The decoder generates text one token at a time, attending back to that representation. It is very close to machine translation, where the source language happens to be audio.
Why Whisper mattered #
OpenAI’s Whisper, released as open weights in 2022, changed the practical situation more than the theoretical one.
It was trained on 680,000 hours of audio, scraped broadly from the web rather than curated in a lab. That is a staggering amount, and it was deliberately messy: multiple languages, accents, background noise, varying recording quality, different microphones.
The consequence is robustness. Systems trained on clean studio recordings fall apart on a phone recording in a café. Whisper was trained on so much varied real-world audio that noisy conditions are part of what it expects rather than an edge case.
It was also multitask. The same model transcribes, translates into English, identifies the language, and detects when there is no speech at all. These were separate systems before. It handled 99 languages in a single model, with quality varying by how much of each language was in the training data.
And it was released openly, in multiple sizes from about 75 MB to a few gigabytes. That is what made on-device transcription possible for anyone building an app, because you can ship the model itself rather than calling somebody’s API.
Why it makes the mistakes it makes #
Every failure mode below follows directly from the architecture.
Homophone confusion. “Their,” “there,” and “they’re” are acoustically identical, so there is no signal in the audio to distinguish them. The model picks based on context, exactly the way a human listener does, and when the context is ambiguous it guesses. This is not a bug that will be fixed, because the information is genuinely not present.
Proper nouns and jargon. Unusual names, technical terms, and company names appear rarely or never in training data, so the model substitutes something phonetically similar and more common. This is why transcription of a specialist conversation is noticeably worse than transcription of small talk.
Invented punctuation. Speech has pauses and intonation, not commas and periods. The model infers punctuation from phrasing and from what text usually looks like. Reasonable guesses, occasionally wrong ones, and entirely a matter of inference rather than recognition.
Hallucination during silence. This one surprises people. Because the decoder is a generative model trained to produce plausible text, feeding it silence or noise can produce confidently generated text that was never spoken. Whisper is known for occasionally emitting things like a subtitle credit line during long silences, because such lines appeared in its training data at the ends of videos. It is the same underlying mechanism as hallucination in language models: the model is built to generate plausible output, and it does that even when the input contains nothing.
Overlapping speakers. Two people talking at once produces overlapping spectrograms and the model transcribes something garbled. Separating speakers, called diarization, is a different task and usually a different model layered on top.
Accents and dialects are recognized in proportion to their presence in the training data. This is a real and well-documented equity problem in the field, not an incidental one.
The size and speed tradeoff #
Whisper ships in tiers, and the choice between them is a genuine trade rather than a formality:
| Model | Size | Character |
|---|---|---|
| Tiny | ~75 MB | Fast, rough. Fine for clear speech and quick notes |
| Base | ~142 MB | Balanced. A reasonable default |
| Small | ~466 MB | Noticeably better on accents and noise |
| Medium | ~1.5 GB | Strong, slower |
| Large | ~3 GB | Best available, needs real hardware |
Bigger models are better at exactly the hard cases: accents, background noise, technical vocabulary, and unusual names. On clean audio from a native speaker in a quiet room, the difference between tiny and large is much smaller than the file sizes suggest.
Why this runs on your phone #
Speech models are small compared to language models. A 466 MB model is unremarkable as an app download, and running it does not strain a modern phone.
Combine that with the fact that audio is one of the more sensitive things you can hand to a server, and on-device transcription becomes an obvious fit. Cloud transcription services upload your recordings, process them on their infrastructure, often retain them, and in some cases use them for training. For a voice memo about groceries that is fine. For a medical appointment, an interview with a source, a therapy session, or a confidential call it is not.
I built Private Transcribe on exactly this premise: Whisper running locally, four model tiers so you can pick your own speed and accuracy point, 99 or more languages, and audio that never touches a network because there is no network code in the transcription path. The engineering write-up covers what was actually hard about it, which was mostly memory management on long recordings rather than the model itself.
The broader argument for this architecture is in on-device AI explained, and the compression that makes it fit is in quantization.
Getting better results #
These apply whatever tool you use.
Get the microphone close. Distance is the biggest single factor in transcription quality, far more than the model tier. A phone at arm’s length beats a laptop across a table.
Reduce background noise rather than compensating for it afterward. Air conditioning, traffic, and café ambience all degrade accuracy meaningfully.
Use a bigger model for hard audio. Accents, jargon, multiple speakers, and poor recordings all benefit disproportionately.
Expect to fix proper nouns. Names of people, companies, and technical terms are where you will spend your editing time. Everything else is usually close.
Record uncompressed if you can. Heavy audio compression removes exactly the high-frequency detail that distinguishes consonants.
The short version #
Audio becomes a spectrogram, which is effectively an image. A transformer reads that image and generates text, one token at a time, the same way a language model generates text. Whisper made it work in the real world by training on an enormous amount of messy audio and then releasing the weights. The errors it makes are the errors that architecture implies: homophones, rare names, invented punctuation, and occasional confident fabrication over silence.
More on the models underneath in how LLMs work, and on why generative models confabulate in why LLMs hallucinate.