Most of the failure cases we saw weren't accent-related at all. They were rooms — HVAC, a projector fan, two people talking over each other.
Interviewer
11:19
And the gating happens before the model, not after?
Reyes
11:22
Before. We run spectral subtraction and an RNN suppressor on the buffer, then hand a clean 16-kilohertz stream to the transcriber. Nothing downstream has to guess.
Reyes · live
12:04
so the noise floor drops about eighteen decibels before the model