Simply Broken

← A Voice Reader for Simply Broken

v2 · the voice

Six takes on the same paragraph

v1 proved the machinery works and the engine is good enough. The voice was the problem. This is the same opening read six ways: the stock voice as a control, then five clones of Amit's own, varying the reference cut and the two settings that actually change how a sentence lands.

“The chatterbox is pretty good. The voice itself though I don't like. Can you make the voice a bit more manly maybe a bit more similar to my voice.”

01 · ListenThe rack

Every take says exactly the same 232 characters, so differences are the voice and nothing else. Playing one stops the others, which is the only way an A/B is worth anything.

Control · stock voice
What v1 shipped. No cloning. This is the one you did not like.
23.5s~210 wpm
A · natural
Cloned from your voice, settings left alone. The straight comparison against the control.
31.9sexag .5 · cfg .5
A · flat authority
Same clone, less animation and a looser pull toward the reference. Reads heavier and more deliberate — the closest thing here to “more manly”.
32.7sexag .3 · cfg .4
A · slower
Same clone, pulled furthest from the reference, which is the documented way to slow a fast speaker down.
30.8sexag .5 · cfg .3
B · different reference cut
Same settings as A natural, but cloned from a later 22 seconds of the same recording. A test of how much the choice of reference matters.
31.4sexag .5 · cfg .5
C · a third reference cut
Later again. Included because reference choice turned out to be the one variable with a measurable failure.
31.1sexag .5 · cfg .5

02 · What changedThree levers, and what each one does

The clone. Chatterbox copies a voice from a short reference with no training step — 22 seconds of clean speech is the useful amount. The reference here is your own recording. It is Hebrew, and the output is English: the timbre carries across languages even though the words do not, which is why this works at all.

Exaggeration is emotional intensity. Lower reads flatter, heavier and more certain; higher is more animated, which for a business essay reads as less serious. “More manly” is mostly this lever, so take A flat authority is the one to judge that on.

cfg weight is how hard generation is pulled toward the reference. Counter-intuitively, lower slows the delivery down. That matters because the control speaks at about 210 words per minute against an audiobook norm near 150.

The pace problem solved itself

Cloning fixed the speed without anyone asking it to. The same words take 23.5 seconds in the stock voice and about 32 seconds cloned — roughly 30% slower, which puts it near 155 words per minute instead of 210. The stock voice was not just wrong in character. It was rushing.

03 · EvidenceEvery take was checked, not just listened to

A cloned voice can drawl or repeat itself, and a longer file looks the same either way from the outside. So each take was transcribed back and compared against the script it was given. Length here is real slower speech, not a defect — but the check found two takes that drop or bend a word.

TakeTranscribes asVerdict
A · naturalword perfectclean
A · flat authorityword perfectclean
A · slower“can't slip” for “can't-sleep”one bent word
B“begins retelling” for “begins with telling”drops a word
CMost follow rules” for “Must follow rules”bends the title

So the reference cut is not a neutral choice. Cut A is the one to keep — both takes built on it came back exact, while both alternative cuts introduced an error in 232 characters. Over a whole article that rate would be felt.

04 · NextWhat happens after you pick

Say which take sounds like you and the full article gets rendered in it, replacing the stock-voice sample on v1. A full render in a cloned voice takes roughly ten minutes of machine time for three minutes of audio.

Two things stay true whichever take wins. Chatterbox stamps an inaudible watermark on every file it makes, unconditionally. And a cloned voice is only free while it runs on this machine — the whole 43-piece corpus is 15 to 20 hours of local rendering against about five dollars on a paid engine, which cannot clone you.

Still open from v1

Scope (all 43 pieces, blog only, or blog plus four explorations), whether images get described in the audio at all, and whether the AirPods buttons skip fifteen seconds or switch between summary and full. None of those depend on the voice.