← A Voice Reader for Simply Broken
v4 · deeperFive voices with more weight
Round one cloned the author and fixed the pace, but not the character. Round two goes after depth and authority directly: the cloning reference itself is pitched and shaped down before the model ever hears it, so the clone inherits a bigger voice rather than just a lower one. Same paragraph as every previous round.
“I still don't like the voice. Generate a new version with five different voice options. Make them more authoritative and deeper bass as one example and another example and try a few others.”
01 · ListenThe rack, deepest first
Two carry-overs sit at the bottom as references: round one's straight clone and the original stock voice. Playing any take stops the others.
02 · The trickTransform the reference, not the output
Chatterbox clones whatever it is handed. So instead of asking the model for a deeper voice — there is no knob for that — the 22-second reference recording is pitched down with its formants riding along, which is what makes a voice read as physically larger rather than merely lower. The model then clones that bigger voice natively, at normal speed, with normal articulation.
Two takes also use a post-render shift, where the whole output is settled down after synthesis. It is the simpler lever, and its cost shows up in the evidence table: the shifted audio slightly blurs the transcription check, which is the canary for what it does to consonants.
03 · EvidenceChecked, not just heard
| Take | Transcribes as | Verdict |
|---|---|---|
| Deep · −2st | word perfect | clean |
| Warm authority | word perfect | clean |
| Deepest · −3st | “Tché, must follow rules for finalizing storytelling” | articulation smears |
| Gravitas | “Sir, must follow” for “3 Must follow” | opening numeral bends |
| Stock deepened | “30. Must follow” | opening numeral bends |
Two semitones of reference shift is free — the clone stays word-perfect. Three is not: the model starts smearing consonants (“fundraising” became “finalizing”). And both post-render shifts bent the spoken “3” at the very top of the piece. If the winner is one of the flagged takes, the fix is to hold its character and back the shift off half a semitone, then re-verify.
04 · NextPick one, or steer
Name a take and the full article renders in it overnight, replacing the sample on v1. Or steer — “deep but warmer”, “the second one but slower” — and the next rack starts from the winner instead of from scratch. Every knob here composes with every other.
Worth restating from v2: whichever voice wins, Chatterbox watermarks every file, and the full 43-piece corpus at local speed is 15–20 hours of machine time against about five dollars on a paid engine — which cannot do the reference-transform trick, but ships proper SSML pacing instead.