Recording speech in a noisy room
Not all noise is equal. Knowing which kind you are fighting tells you what to do about it.
· 6 min read
Every transcription tool performs beautifully in a quiet room. What separates them is what happens in a real one — a café, an open-plan office, a hall with a projector fan.
It helps to know that "noise" is not one problem. It is at least three, and they need different responses.
Steady noise is the easy kind
Air conditioning, a fan, traffic hum, a fridge compressor: these are continuous and broadly unchanging. That predictability is exactly what makes them tractable — a system can characterise a constant background and separate it from speech, because speech is the part that keeps changing.
Steady noise sounds bad to a human ear and is often the least damaging to a transcript.
Other people talking is the hard kind
Background speech is the worst case, because it has exactly the same statistical character as the speech you want. There is no acoustic signature that separates "the voice I care about" from "a voice two tables away" — both are human speech in the same frequency range with the same rhythms.
This is why a busy café is harder than a building site, even though the building site is louder. If you can move away from other conversations, do that first.
Reverberation is the invisible kind
A hard, empty room sends every sound back at the microphone a few milliseconds late. The result smears the boundaries between words — consonants in particular — and is the reason a recording can sound "fine" to you in the room and come out mushy.
Reverberation gets worse with distance, which is another reason to close the gap between mouth and microphone. Soft furnishings help; empty glass-walled meeting rooms are the worst offenders.
What actually helps, in order
Ordered by how much difference they make per unit of effort:
- Move the microphone closer — halving the distance is worth more than any software setting
- Move away from other conversations, even by a few metres
- Put something soft in the room, or pick a room that already has some
- Turn off what you can: fans, air conditioning, notifications
- Point the microphone at the speaker and away from the noise source
What to expect from the software
Modern speech recognition is trained on realistic audio, not studio recordings, so it degrades gracefully rather than falling off a cliff. Expect accuracy to drop with noise rather than collapse — and expect the errors to land on names, numbers and technical terms first, because those are the words that cannot be reconstructed from context.
That last point is worth planning around: if a recording is going to be noisy, decide in advance which details you will verify afterwards.