→ all versions of this pipeline
You asked whether it is really this hard. The honest answer is no — and the reason it stayed audible is that I was cutting far too gently.
What I measured after your last round. Those spans are 93%, 97% and 100% noise with no speech in them at all. I had built careful machinery to protect a voice that is barely present, and then let that caution set the cut. Bringing the noise down to the room needed 10 to 21 dB; I applied 7 to 13, and then smoothed the gain over a quarter of a second, which averaged much of that away again.
| last time | now | |
|---|---|---|
| cut applied to the noise | 7–13 dB | 8–29 dB |
| where the noise ends up | still well above the room | at or below the room |
| how the gain is sized | one figure per span, from a median | per 30 ms block, aimed at the room |
| ramp | 0.25 s across everything | 0.15 s at the edges only |
Two arithmetic mistakes underneath it: I sized the cut from the median of the noise, which sits well below how loud it actually sounds, and I smoothed the gain across the whole span so that any protected moment averaged the cut away around it.
Each span now has three buttons — the original, what you heard last time, and this pass. The middle one matters: if this sounds no different from last time, then my measurements are lying and I would rather know that immediately.
Two new things to listen for. Hollow — some spans now sit below the room, so the noise may be replaced by a hole. And speech damaged, since the cuts are much deeper than before.
Three buttons: the original, what you heard last time, and this pass. The middle one is there so you can judge whether this is actually different or just my arithmetic.
Same three. On this episode the boundaries are still mine, so if a span sounds untouched it is the boundary, not the cut.