Your approach works — ten of eleven, and it wins on every span where
everything else failed. Before it goes anywhere near the archive, two things have to be
checked, and neither is about the rustle.
Why this now matters more than the noise. Keeping the voice and rebuilding the
room is applied to the whole episode, not to marked spans. That is why it beat
everything else — and it means the entire localisation problem disappears. Six failed
attempts, three rounds of your listening and spans that were off by two seconds were all
in service of finding boundaries this approach does not need.
But it also means the decision is no longer "treat these three seconds". It is
discard everything in the episode that a model does not consider voice. That is a
much larger claim, and it has never been checked.
The risk. In a bedtime story with two children, the discarded channel may hold a
laugh, a kiss, a clap, a toy, a page turning — things that are the point of the
recording rather than defects. If any of that is in there, we are trading a rustle for
something irreplaceable.
So below is the discarded channel itself, lifted so it is audible, at its twelve
loudest moments in each episode. Each one pairs with the original at the same instant so
you can hear what was there and what got taken.
First: does it still sound like them?
The other gate, and the one your standing rule is about. Short A/B
clips cannot answer it — voice quality shows itself over minutes. Play a few of these from
the start and listen to Hili and Rani rather than to the noise.
nothing playing
Episode 6 — the twelve loudest discarded moments
The red button is the discarded channel, amplified. The green one
is the original at the same moment for comparison.