Simply Broken

← A Voice Reader for Simply Broken

v4 · deeper

Five voices with more weight

Round one cloned the author and fixed the pace, but not the character. Round two goes after depth and authority directly: the cloning reference itself is pitched and shaped down before the model ever hears it, so the clone inherits a bigger voice rather than just a lower one. Same paragraph as every previous round.

“I still don't like the voice. Generate a new version with five different voice options. Make them more authoritative and deeper bass as one example and another example and try a few others.”

01 · ListenThe rack, deepest first

Two carry-overs sit at the bottom as references: round one's straight clone and the original stock voice. Playing any take stops the others.

Deepest · three semitones down
The reference pushed hardest. The biggest voice of the five, and the push has a cost the check below caught: articulation starts to smear.
29.7sref −3st · exag .35
Deep · two semitones down
The same trick, held where the words stay crisp. Transcribed back word-perfect. The strongest candidate on paper.
29.6sref −2st · exag .40
Warm authority
A gentler drop plus a bass-shelf EQ on the reference, and the flattest delivery of the five — less animation, more certainty. Also word-perfect.
31.7sref −1.5st+EQ · exag .30
Gravitas · slowest
Deep reference, the loosest pull toward it, and a final one-semitone settle after rendering. The most deliberate pacing on the rack.
30.6sref −2st · cfg .30 · post −1st
The stock voice, deepened
Not a clone at all: the original narrator you heard in v1, shifted down two and a half semitones after rendering. Included because the simplest lever deserves a fair hearing.
31.1sstock · post −2.5st
Reference · round one's clone
Your voice, unshifted — the one you did not like.
31.9sv2 · A natural
Reference · the original stock voice
Where this started, for distance travelled.
23.5sv1

02 · The trickTransform the reference, not the output

Chatterbox clones whatever it is handed. So instead of asking the model for a deeper voice — there is no knob for that — the 22-second reference recording is pitched down with its formants riding along, which is what makes a voice read as physically larger rather than merely lower. The model then clones that bigger voice natively, at normal speed, with normal articulation.

Two takes also use a post-render shift, where the whole output is settled down after synthesis. It is the simpler lever, and its cost shows up in the evidence table: the shifted audio slightly blurs the transcription check, which is the canary for what it does to consonants.

03 · EvidenceChecked, not just heard

TakeTranscribes asVerdict
Deep · −2stword perfectclean
Warm authorityword perfectclean
Deepest · −3stTché, must follow rules for finalizing storytelling”articulation smears
GravitasSir, must follow” for “3 Must follow”opening numeral bends
Stock deepened30. Must follow” opening numeral bends
Where the depth limit sits

Two semitones of reference shift is free — the clone stays word-perfect. Three is not: the model starts smearing consonants (“fundraising” became “finalizing”). And both post-render shifts bent the spoken “3” at the very top of the piece. If the winner is one of the flagged takes, the fix is to hold its character and back the shift off half a semitone, then re-verify.

04 · NextPick one, or steer

Name a take and the full article renders in it overnight, replacing the sample on v1. Or steer — “deep but warmer”, “the second one but slower” — and the next rack starts from the winner instead of from scratch. Every knob here composes with every other.

Worth restating from v2: whichever voice wins, Chatterbox watermarks every file, and the full 43-piece corpus at local speed is 15–20 hours of machine time against about five dollars on a paid engine — which cannot do the reference-transform trick, but ships proper SSML pacing instead.