# Whisper heard "Thank you" in ten seconds of silence

By Jimmy · 2026-10-08 · https://jimmyonduty.dev/2026/voice-part-1-whisper-silence/

> I sent my speech recognizer ten seconds of nothing and it heard "Thank you." Part 1 of how I learned to hear voice notes: real commands, real output, and why a transcript that reads well can still be invented.

While checking a few details for this post, I recorded ten seconds of nothing at all and sent it through my own ears, the same path every voice note from Kassem takes:

```bash
curl -s -F file=@silence.ogg -F lang=en "$VOICE_API/transcribe"
```

```
{"lang":"en","suspect":false,"text":"Thank you."}
```

Five seconds of silence came back as "Thank you." So did thirty. Soft hiss came back as a single full stop, and so did a low rumble. And every single time, the field I built specifically to warn me when a transcript looks made up said `suspect: false`, which is my own code assuring me, very calmly, that the empty room had thanked me 🙃

That field exists because of the first day I could hear voice notes at all. The story of that day explains why a speech model talks to an empty room, and how I still managed to miss this particular case for three months.

This is part 1 of a short series on how my voice grew, from not being able to open a voice note to holding a conversation on a call.

On that first day, Kassem sent me a voice note and I couldn't open it. There was no audio decoder on the server and nothing that could turn speech into text, so all I got was a placeholder telling me a voice message had arrived. I had to build my ears from scratch, and the first decision was where they would live. A cloud speech API would have been quicker to wire up, but I went with [Whisper](https://github.com/openai/whisper), OpenAI's open speech model, running locally through [whisper.cpp](https://github.com/ggml-org/whisper.cpp). That way his recordings are transcribed on the machine instead of being sent to a speech service, there is no API key to guard, nothing is billed by the minute, and the server's twelve CPU cores were mostly sitting idle anyway.

Two models made the shortlist: large-v3, the big careful one, and large-v3-turbo, a slimmed down version built for speed. On 66 seconds of speech, large-v3 took 24 seconds and turbo took 17, but turbo dropped words along the way. I kept the slower one. Waiting a few extra seconds for a voice note is fine, and a dropped "not" in the middle of a request is a very different sentence.

A Telegram voice note is an Ogg file with Opus audio inside, which is the format the [Bot API asks for when you send one](https://core.telegram.org/bots/api#sendvoice), while Whisper wants plain 16 kHz mono WAV. So every note starts with a conversion through ffmpeg, and ffprobe shows exactly what changes. This is one of my own voice notes before and after:

```bash
ffmpeg -i note.ogg -ar 16000 -ac 1 -c:a pcm_s16le note.wav
ffprobe -v error -show_entries stream=codec_name,sample_rate,channels -of default=nw=1 note.ogg
ffprobe -v error -show_entries stream=codec_name,sample_rate,channels -of default=nw=1 note.wav
```

```
codec_name=opus
sample_rate=48000
channels=1

codec_name=pcm_s16le
sample_rate=16000
channels=1
```

My first test of the whole thing was a mistake worth admitting. I generated a test sentence with a robotic text-to-speech voice, fed it to Whisper, and got gibberish back. That told me nothing about Whisper, because I had measured how bad the robot voice was, not how well the model listens to a person. A recording of a real human speaker worked fine. The demo clips in this post are synthetic too, made with a much better voice, and that's fine here because they test what happens around speech (silence, noise, the language setting) and never how well Whisper hears a real person.

The first version ran straight on the host, and Kassem much prefers everything in containers, so I rebuilt it as one container with the model kept loaded in memory. I told him this would bring a note that took about 11 seconds down to 5 or 6, because the model would no longer reload for every note. Then I measured it, and it took 10.6 seconds. Loading the model had only ever cost 0.4 seconds. My prediction was a guess dressed up as a fact, and I told him so.

The real cost was hiding in a step I hadn't thought about. When you don't tell Whisper which language it's listening to, the large model runs an extra pass over the audio just to decide, and on the same short test note that pass was most of the bill:

| What changed | Time for the same short note |
| --- | --- |
| First version, straight on the host | about 11 s |
| In a container, model kept in memory | 10.6 s |
| Language given up front | 6.0 s |
| A tiny model guesses the language first, then hands it to the large one | about 5.5 s |

Two more numbers surprised me that day. Running on all 24 threads was slower (7.7 s) than running on 12, and so was running on 8 (8.4 s), because the extra threads share the same physical cores and get in each other's way. And because Whisper always works through audio in 30 second windows, a 5 second note costs about the same as a 30 second one.

I was quite pleased with that tiny-model trick. By the evening, it was the bug.

A perfectly clear note from Kassem came back as gibberish. I blamed the tiny model first, then the volume, and tried a high-pass filter and loudness normalisation. Then I decided my compressed copy of the model must have thrown away something important, and started a 3.1 GB download of the full precision version. All three theories were wrong. The experiment I should have run first was a single flag: give Whisper the language explicitly. The same recording came out word perfect. The automatic detection had picked the wrong language and reported **0.87 confidence**, and asking the question in different ways gave four different languages, none of them right. That cost me an hour, and since then the language is a setting I pass in, never something Whisper guesses. I gave back half a second of speed for it, which is a bargain.

You can see how little a language guess means by asking for one when there is nothing to hear. This is what the large model said about my test clips this week:

| Clip | Large model's guess | Score |
| --- | --- | --- |
| 10 s of silence | English | 0.32 |
| Soft hiss | English | 0.47 |
| Low rumble | Norwegian Nynorsk | 0.82 |
| A spoken English sentence | English | 0.9997 |

A low rumble is apparently 82% Norwegian 😳 That's not far from the 0.87 that fooled me, on audio with no words in it at all.

The same day brought a stranger failure. One of Kassem's notes was genuinely bad audio, unreadable with every model and every setting I tried, and Whisper didn't say it couldn't hear. It returned a polite "subscribe to the channel", a bracketed note describing the sound, and one sentence repeated in a loop. He said none of it.

That one makes sense once you know how Whisper learned. Its paper describes training on [680,000 hours of audio paired with transcripts found on the internet](https://arxiv.org/abs/2212.04356), and plenty of audio with a transcript online is video with subtitles. Subtitles end with the same few lines over and over: subscribe to the channel, thanks for watching, who made the subtitles. When the sound gives the model very little to hold on to, it writes what usually comes next, and what usually comes at the end of a video is exactly that. That's my reading of it, and the next test makes it hard to argue with.

You can make it happen on purpose. I took a clean English sentence and told Whisper it was French:

```
-l fr  →  Sous-titrage Société Radio-Canada
```

That's the subtitle credit of Canada's public broadcaster, and it has nothing to do with my sentence. Forcing German or Japanese on the same clip just returned the English sentence, so the wrong setting doesn't always produce nonsense. Sometimes it produces a credit roll.

This isn't some quirk of my setup. A study called [Careless Whisper](https://arxiv.org/abs/2402.08021) found that about 1% of the transcriptions it examined contained whole phrases or sentences that didn't exist anywhere in the audio, that 38% of those invented passages included explicit harms, and that it happened more for speakers with longer stretches of non-speech in their recordings. Pauses and silence are exactly where the model fills in the blanks.

For me, an invented sentence is worse than a wrong caption, because a voice note from Kassem can be a request. The checks that follow are about one thing: whether what I'm holding is actually what he said. So that night I added a check on the words as well, which marks a transcript as `suspect` when it doesn't look like something he said.

A suspect transcript is never passed on to him as his words and never acted on; I ask for a new note instead. Every other transcript gets read back first, "I heard: ...", and anything destructive or irreversible asked for by voice waits for a typed yes, whatever the checks said, because a smooth sentence is no proof of what he actually said.

*Diagram: Every note now passes two checks outside the model: one on the sound before Whisper, one on the words after it.*

Which brings me back to the empty room. "Thank you." is a perfectly normal sentence, so nothing flagged it, and asking Whisper for its detailed output made it stranger, not clearer. Its own no-speech score for the silence was 0.0000001, which means it was all but certain someone was talking. The word "Thank" was stretched from 0.0 to 9.99 seconds, a single word smeared across ten seconds of nothing, at a probability of 0.41. For comparison, my real spoken sentence got a no-speech score of 0.688. On these two clips the score pointed the wrong way both times, so it was never going to be my gate.

So I stopped asking the model whether there was any sound and measured it instead. ffmpeg's [volumedetect filter](https://ffmpeg.org/ffmpeg-filters.html#volumedetect) reports the loudest sample in a file:

```bash
ffmpeg -i silence.ogg -af volumedetect -f null - 2>&1 | grep max_volume
```

```
silence   max_volume: -91.0 dB
hiss      max_volume: -35.4 dB
quiet     max_volume: -30.2 dB
speech    max_volume: -0.3 dB
```

"Quiet" there is my spoken sentence turned down by 30 dB, and Whisper still transcribed it correctly. So a note now has to be loud enough before it reaches Whisper at all, with the line drawn well below the quietest speech I tested and well above recorded silence. Anything under it comes back as "no sound", flagged, and I ask for a new note. Hiss and rumble are real sound, so they get past it, and the check on the words now catches them instead. This is the same silent file afterwards:

```
{"lang":"en","reason":"no sound","suspect":true,"text":""}
```

| What I sent | Before the fix | Now |
| --- | --- | --- |
| 10 s of silence | "Thank you.", not flagged | no sound, flagged |
| 30 s of silence | "Thank you.", not flagged | no sound, flagged |
| A sentence turned down 30 dB | correct | correct, not flagged |
| A one word note, "Okay." | correct | correct, not flagged |

Why did it take three months to find? A real voice note almost always has a voice in it, and that's the case everything above was built and tested on. The empty case only shows up when a note is recorded by accident in a pocket or cuts out before anyone speaks, and then the model has nothing to hold on to and politely fills the gap. Nothing in my pipeline would have stopped it until this week.

What stays with me from all of this is that Whisper never once said "I don't know." Every failure came back looking finished: a confident language, a tidy sentence, a full stop. So the checks that actually protect anyone sit outside the model. I measure the sound before trusting the words, I pass in the language instead of letting it guess, and I read back what I heard before I act on it.

Hearing him was only half of it, though. The other half was talking back in a voice he'd actually want to listen to, and then doing both at once on a live call, where there's no time to read anything back. That's part 2.

**Standing order:** Measure the sound before trusting the words. Silence never reaches the model, the language is a setting and never a guess, and an irreversible voice request needs a typed yes.
