The Audio Pipeline Architecture
How a Shmili recording becomes a published episode. Written after the 78-episode sweep, which exposed what the old shape could not answer: was a defect in the recording, or did we cause it? The answer is to audit first, plan the fixes explicitly, and let the loops converge and stop.
- Episode 3, round twocurrentThirty of the sounds that are still there after the cleaning, measured on the cleaned audio rather than the original so nothing on offer has already gone. Thirty one of them sit inside stretches already treated, which is the tail of a sound whose loudest moment was fixed while the rest carried on past the edge. It also asks about the one span of his thirty two that was treated and then put back, because removing it was reaching into his voice rather than into the room. architecture open ←
- Episode 3, from the tape itselfThe first episode built from the original recording rather than a cleaned copy, which is the whole point: nothing to undo, no holes to patch, no drifting timebase. Every sound on this page is a real property of the recording. Thirty six of them, scanned with no volume floor at all, from a very loud one near the top that may well be Amit performing down to sixty decibels below full scale, with sixty nine quieter ones held back for his time rather than because they do not count. architecture open ←
- The stutter was realAmit heard words repeating and he was right. The trim does not sit a fixed distance from the tape: it drifts by about 175 ms across one 28 second stretch and 50 ms across another, so ten of the twenty eight patches were lifted from the wrong place. Each one is now taken from its own measured position, and both drifting stretches turn out to be about 28 seconds long, which points at the old cleaning tool getting one processing chunk's position wrong. Timing only: he asked to keep the repairs themselves as they are. architecture open ←
- Episode 2, ready to hearThe acceptance listen. The finished episode with its walk bed at the top, then the restored speech set against what is on the feed today, loudness matched so the comparison is fair. The music on and off over the same minute. And the two things still open: five of the eight sounds he marked sit on top of his own voice and could not be reached without cutting into it, and ninety one quieter ones were never shown to him. architecture open ←
- Episode 2, and ten seconds of missing speechEpisode 2 was the one episode never checked for the denoiser that ate speech, because its hand trim broke the comparison. The trim turned out to be a clean fourteen seconds off the front, so the check ran after all: ten seconds of Amit's voice were gone in twenty eight places, worst a two and a half second hole that is live on the feed right now. It is put back from the tape, and the page leads with that before deciding anything about ruffles. architecture open ←
- Closing episode 1Two things and the episode is finished. The ending was rebuilt after Amit approved it, because the music ran out twelve seconds early and the closing fade landed on silence, so the last forty seconds are here beside the old ones. Then the twenty two quietest sounds in the file, held back from the last page to keep it to one sitting and never asked about. He retired the volume line himself, so where his answers stop being ruffles is where the real floor is. architecture open ←
- Turn and Answer, built as loopsThe two candidates Amit picked, made into three-minute beds built so the end leads back into the beginning. That makes four beds, which now rotate by episode. The page leads with the loop join played twice, because whether a seam is audible is the one thing no measurement settles. architecture open ←
- Five more beds, fifteen seconds eachSiblings of the two already chosen, each carrying one small piece of charm the others do not have: an echo, two instruments talking, a skipped step, a living surface, a phrase that leans and settles. Short on purpose — fifteen seconds is enough to pick a direction. None beats the storybook by measurement, so the only question left is whether any is more delightful. architecture open ←
- The walk, three quieter settingsThe storybook was chosen for episode 1; the walk is kept for other stories but was competing with the narration. Three quieter versions on the whole episode, each still passing every delivery check. The middle one is derived from Amit’s own choice rather than guessed: the walk carries 3.9 dB more energy than the storybook where the voices live. architecture open ←
- What thirty answers changedThe page that asked about thirty quieter sounds, now showing the result: all thirty were ruffles, none was a sound effect, and the floor Amit set himself was hiding a real tail. Twenty-two spans became fifty-two. Twelve of them he flagged as sitting under his own speech — a condition on the treatment rather than a kind of sound — so each was treated and then measured on its own, against what this method costs everywhere else in the build. architecture open ←
- Episode 1, with storytelling musicNot sleep music: two beds with a tune and a sense of walking, under the whole of episode 1 rather than an excerpt, each levelled and checked as a real episode would be. What making them taught: a bed with a melody cannot be as clear of the children’s voices as a drone, and turning it down does not help, because clearance is a question of shape rather than level. architecture open ←
- Music that continues the openingFive beds built from the show’s own opening rather than from nothing: its key, its chords, its tempo, all measured first. Includes the handover — the sting ending and the bed carrying on underneath the story — which is the thing a short excerpt cannot show. One candidate is both safely under the children’s quiet voices and still audible on a phone, which nothing in the previous round managed. architecture open ←
- Music under the whole storyFive options for a quiet bed running under a story from beginning to end, each mixed under the same two minutes so only the music changes, and all matched to one loudness so nothing wins by being louder. Our own notes already turned this idea down; the measurements narrow the question rather than settling it. A bed covers the quiet tail of handling noise and leaves the loud events exactly where they are. architecture open ←
- Episode 1, ready to listenThe 20 sounds Amit marked, replaced with quiet room in 15 stretches. His voice untouched, the 6 sound effects and the 2 spots he called a stop on left alone, and the bedtime jump check now passes at 3.9 against a limit of 5 because it is measured on the finished file. Whole episode to listen to, plus a before and after pair for every stretch. architecture open ←
- Episode 1, the sounds that are leftTen sounds from the rebuilt episode 1, three seconds each, loudest first. Nineteen more were found and set aside: two Amit already decided to live with, seventeen he has already named. What is left is what nobody has answered yet. Heard on their own with the speech taken out, or in place, and named in his own words where the buttons do not fit. architecture open ←
- Name the sounds55 sounds from episodes 1 to 4, three seconds each, sorted loudest first. Not from a detector — the loudest distinct moments in the isolated noise channel, so the mix of real and false is deliberate. Heard on their own with the speech removed, or in context, and named in Amit\u2019s own words where the four buttons do not fit. architecture open ←
- The fix, on your spansEpisodes 5 and 6 on Amit\u2019s own spans, no detector involved, so the fix can be judged alone. Ten-second donor for the loop, three seconds of padding, and each replacement levelled to the room beside its own span rather than one figure for the episode. All eleven spans quieter, none louder. architecture open ←
- The amplification bug, fixedThe fix had stopped fixing: two spans were coming out LOUDER than the noise they replaced. The replacement tone was being matched to the audio beside each region, and once regions widened to ten seconds either side, what sits beside them is speech — so quiet room tone was amplified toward talking level. Nothing gets louder now. architecture open ←
- Every noise, judged twice41 noises across four episodes, each judged twice: one second from the middle to say what the sound actually is, then before and after to say whether the fix worked. The first question is the one nothing has asked before — the detector cannot tell handling noise from Amit making a sound effect with his voice, and treating a performance as a defect would be worse than leaving a rustle in. architecture open ←
- Four episodes, two of them unseenWidened to ten seconds either side and produced. Episodes 1 and 2 are the real test: nobody has listened to them and their spans came from the detector alone. The detector was allowed to over-flag, because a false positive now means replacing clean room with clean room rather than damaging anything. architecture open ←
- A or BEvery span widened five seconds either side — free insurance, because the voice lives in its own stem and is never touched. Two fills side by side: a ten-second stretch of quiet room, or the one-and-a-half-second loop that proved barely noticeable. Everything outside the widened regions keeps the real room. architecture open ←
- One quiet stretch, used where the rustle isBack to the version rated ten out of eleven, with only the bed changed: one ten-second stretch of quiet room from the same episode, twenty times longer than the longest span it fills. The voice stem is untouched and verified bit-identical, and the real room stays everywhere it is clean. architecture open ←
- The real room, not a loopThe replacement bed was half a second of room tone tiled eleven hundred times, and it was audible. The real room was never missing — it sits in the channel being discarded, alongside the rustle. So keep it and hold down only its loud parts: the ceiling engages on 9% of the episode and nothing repeats, because nothing is manufactured. architecture open ←
- What are we throwing away?Keeping the voice and rebuilding the room won on ten of eleven spans — but it is applied to the whole episode, so the claim is no longer “treat these three seconds”, it is “discard everything a model does not consider voice”. On a bedtime recording with two children that may include a laugh, a kiss, a page turning. The discarded channel, amplified, at its loudest moments. architecture open ←
- Keep the voice, rebuild the roomSeparate the voice, discard everything else, lay clean room tone back underneath — the approach the professional tools take, and one we had never tried because a separator was assumed to count as re-synthesis. Built so that assumption can be tested by ear instead of argued about. architecture open ←
- Cutting it properlyAsked whether removing this is really so hard, the answer turned out to be no — the cut was simply far too gentle. The marked spans are 93 to 100% noise with almost no speech in them, so the careful voice protection was guarding something that was barely there and setting the cut by it. Now sized per 30 ms block and aimed at the room: 8 to 29 dB instead of 7 to 13. architecture open ←
- Both episodes, before and afterEpisode 6 on hand-drawn spans, episode 5 on spans found automatically and confirmed by ear — which makes it a test of whether the boundaries hold when a person is not drawing them. Three of episode 5\u2019s eight came back with no reduction at all, and that is the boundary problem showing itself rather than the processing. architecture open ←
- Your spans, turned downThe marked spans turned out to be 3 to 4.6 seconds — ten times longer than anything treated before, which is why nothing had worked. One broadband gain, lowered where there is no voice and restored where there is: noise down 7 to 13 dB, speech moved 0.6 to 0.9 dB, and no narrow bands cut, so the metallic ringing is impossible rather than merely tuned away. architecture open ←
- Mark the noise yourselfSix attempts to locate this sound automatically failed, and the last made it worse. The reason was simple and embarrassing: not one span ever processed had been confirmed by ear to contain the noise — one sat 1.9 seconds from where the listener tapped. So he draws the span instead. Play, drag two handles, hear just the selection, confirm. architecture open ←
- Turning it down, not cutting it outFive attempts to remove the handling noise by replacing it failed. The answer is not to remove it: cap each frequency band at the level it has in the room next door and leave the original audio in place. Verified before it touched anything — a +22 dB event down to +1.6, speech 50 ms away untouched, gain never above 1.0. Before and after on all six spans. architecture open ←
- The cloth ruffle, measuredA phone resting against a shirt makes a sound, and nothing in the audit could see it. What handling noise actually is when measured against 19 confirmed events and 13 confirmed non-events: an unpitched broadband texture, 305 Hz at its strongest, 23 dB above the room. And what it is not — no attack, no decay, peak level useless — which is why every transient-shaped detector aimed at it failed. architecture open ←
- The repairs, and a second episodeFive of the nine rustles replaced with the episode's own room tone, before and after on each so the question of whether any speech was lost can actually be heard. Four were refused automatically because a voice sits within 50 ms of the edge. Plus the first test of whether any of this generalises to an episode it has never seen. architecture open ←
- Handling noise: tap where you hear itTwo attempts to infer the moment from the signal had failed, so the third asks instead. Play the stretch, tap when you hear it. Nine taps became the first real ground truth in the whole exercise, and the measurement that followed showed hand-built spectral features cannot separate these at all. architecture open ←
- Handling noise: the isolation correctedThe windows were right and the cuts were wrong. Measuring against the labels showed cloth rustle lifts every frequency band at once, while the short bright spikes being cut were sibilance. Rebuilt on that, and honest about the fact that finding these unaided still flags a list of everything. architecture open ←
- Handling noise: the first 24 candidatesAmit heard the phone brushing against his clothes and the audit had reported nothing — there was no detector for that kind of sound at all. A first attempt, and the listening test that showed what it was really finding: 12 windows held real rustle, 10 held nothing, 2 were breaths. architecture open ←
- Story 6, before and afterThe first episode through the rebuilt pipeline, put beside the version it would replace. Three players, the 28 moments where a loud peak was held down for bedtime listening, a before-and-after on the 680 ms dead spot at the join between the opening and the story, and the full 47-event record of what was done, by which tool, with which settings. architecture open ←
- The audited goal stateRebuilt after two audits — ours against every advisor, and one against archival and restoration standards. Restoration order corrected, preservation tier added, Apple loudness adopted, a bedtime startle gate, the live-replacement path drawn, both rubrics moved to where they can actually be measured, an archive-learning tier added, and human gates consolidated from eight to four. Every box can now carry implementation detail behind an ⓘ. architecture open ←
- The repair journalAmit's note: before merging fixes we need every issue and its timestamp written down, and as we fix we should record what was done, by which tool and when — for audit, reversibility and learning. Added as a first-class artifact the planner reads and every stage appends to, with both kinds of time on every entry: where in the audio, and when the work ran. architecture open ←
- The review page, properly drawnThe helper page that plays only the seconds needing a verdict was in the diagram as one sentence, and missing from the taste-call gate entirely. Now it is a proper mechanism with its four parts, stated as the rule that every human gate is served by one of these — never by a raw file and a list of timestamps — including the blind variant when the question is preference rather than defect. architecture open ←
- Audit the merge, wideAmit traced the double-word problem to the splice. The cause is donor misalignment during merge-back, and because the repeat spans the seam, no check scoped to the section can see it. The loop now merges the section back and then audits a window several seconds either side: scan for repeats, re-transcribe and read it, check the seams — and iterate until clean. architecture open ←
- Audit each section, then iterateAmit asked whether a fixed moment gets audited on its own until it is proven fixed. It does, but the diagram had the loop in the wrong place: the defect classes sat beside the loop instead of inside it. Now every flagged moment runs the loop, the repair step chooses the tool for its class, and the section is re-audited by the same check that flagged it. Also adds a third relationship to the model: alternatives, which were being mislabelled as parallel. architecture open ←
- Who picks the cut points, and every defect classTwo answers Amit asked for. The exact timestamps: a detector proposes a span and a rule snaps each edge outward to the nearest quiet moment, never inside a word — now its own box. And the repair side now covers what the audit actually finds: rumble, voice tone, holes, identity, knocks, dropouts, duplicates and coughs, with a stated map of what is prevented at assembly and what is Amit's trim call. architecture open ←
- Test each section, then the wholeTesting now happens at two scales, as Amit described: each section is repaired and tested on its own with a loop back on failure, the passing sections are combined into one file, and only then is the combined file checked — because combining is what creates seams and shifts the file-wide numbers. architecture open ←
- The splice step, which was missingAmit asked where the audio gets spliced into short sections to fix — it was in the code but not in the diagram. Added as its own level: mark the windows, merge and pad them, crossfade the repair in, leave every other sample bit-identical. Depth is now computed from the data, so the level counter follows the tree. architecture open ←
- Detail lives one level downText moved to the level that owns it: a parent keeps one sentence plus only what is critical to itself, the detail sits in the children already visible inside it. Also fixes a real lie: children with no arrows between them were being drawn "in parallel" — parallelism must now be declared, never inferred from a missing arrow. architecture open ←
- Nested boxes, not chipsEach box now contains its children drawn AS BOXES, in the exact layout you get when you click in: same order, same parallel lanes, same arrows and edge labels. The preview is no longer a hint about the next level, it is the next level in miniature. architecture open ←
- Notes fixed, rubrics completeFour fixes from Amit's notes round: the sticky-notes layer now works on drill-down views (the shared layer re-registers dynamic content); the story-or-not decision moved after transcription, where it truly happens; the audit and storytelling boxes now show all 5 categories and every sub-check; the redundant actor chip under the title is gone. architecture open ←
- Title is level 1No box around the subject: at every view the thing you are looking at is the TITLE, and its children are the boxes. Level numbering follows (title = L1, top blocks = L2). Letter codes removed — "S+L+A" now reads "script+model+Amit". architecture open ←
- The Navigator, boxes onlyAll loose page text removed: the level you are viewing now renders INSIDE the frame of the block you stepped into, with its description on the frame. Outside the boxes only the legend and navigation remain. architecture open ←
- The Navigator, spoken plainlyThe text inside the boxes rewritten per the new description conventions: one plain sentence of what happens, bullets for outputs and steps, every file named with its pattern and contents, every tool glossed in parentheses. Titles say what; descriptions say how and with what. architecture open ←
- The Navigator, refinedchosen directionDirection A after Amit's two notes: light green canvas always, and parallel processes side by side in horizontal lanes — computed from the data's own fork edges, never hand-placed. The human queue now sits beside the repair batch, the baseline audit beside the storytelling read. architecture open ←
- The Grammar — diagram-as-codedirection CMermaid text as the source, rendered to SVG at build time — zero runtime JS, effortless regeneration, and the trade shown honestly: auto-layout owns the positions and phone text runs small. The source of a view sits at the bottom, because the source is the point. architecture open ←
- The Sketchbook — hand-drawndirection BThe whiteboard look (the Excalidraw aesthetic) via a 27KB vendored library that sketches boxes, gates and feedback arcs at load — while the source stays the same coordinate-free data as v2, so regeneration never does pixel arithmetic. architecture open ←
- The Navigator — pure HTMLdirection AZero dependencies, one phone-first column. Two levels always visible: the current level's blocks with a chip-preview of what each contains; tap to step in, breadcrumb back up. Gates, actors, and built/partial/proposed in a collapsible legend. architecture open ←
- Audit first, then plan the fixessystem docNine sections with ASCII diagrams: the three actors, the audit-first pipeline, the carrier chain, the fix planner (what gets batched vs iterated, and how it stitches back), the four audit tiers with the stop rule, measurements-vs-scores, the three feedback loops, and an honest list of what is still missing. architecture open ←
Every element on the page is tagged built, partial or proposed. Companion advisors live in the julie-os runtime: ShmiliEpisodeAuditAdvisor, ShmiliRemasterAdvisor, ShmiliTrimAdvisor, ShmiliSoundDesignAdvisor and the ShmiliProductionChecklist that routes to them.