Simply Broken

← A Voice Reader for Simply Broken

v6 · chasing brian

Can free get to Brian?

Brian won v5. This round chases him with two local engines: Chatterbox cloning Brian's own clip, and Kokoro — a second engine, researched and installed today — whose native American narrator presets sidestep the accent problem entirely. Brian himself sits at the top of the rack as the target.

“Brian is great. Let's see if we can get to it locally. The other voices you generated locally sound Indian accent which is not good for my needs.”

The accent, diagnosed

Chatterbox clones its REFERENCE. Every earlier local take was cloned from your Hebrew recording, so the clones carried a non-native English accent — that is what read as Indian. The fix is in the reference: clone from Brian's clean American clip, or use an engine whose voices are native American English to begin with. Both are on this rack.

01 · ListenThe target, then the chasers

The target
Brian · ElevenLabs
The v5 winner. Everything below is trying to be this, for free.
25.8spaid
Chatterbox, cloning Brian's clip
Brian, cloned
The straight attempt: Brian's 22 seconds as the reference, default settings. Word-perfect on the check.
27.2slocal · $0
Brian, cloned · slower and flatter
Same clone, less animation, looser pull — the settings that read as weight in earlier rounds. Also word-perfect.
27.6slocal · $0
Brian, cloned · bass shelf
The reference EQ'd warmer before cloning. The check caught it bending the opening (“Thermust”) — listen knowing that.
26.1slocal · $0
Kokoro — the second engine, installed today
Onyx
Kokoro's deepest American male. Not a clone — a curated native-English preset, which is why there is no accent to inherit.
27.3slocal · $0 · 6x realtime
Fenrir
Darker and heavier than Onyx — the gravitas end of Kokoro's roster.
28.3slocal · $0 · 6x realtime
Michael
Warmer and more conversational — closer to Brian's comfort than his depth.
31.8slocal · $0 · 6x realtime

02 · The second engineWhy Kokoro, and what it changes

The research pass ranked the 2026 open-model field against three hard requirements: commercial licence, Apple Silicon, top-tier English. Most of the acclaimed names fell at a gate — Fish Speech's weights are non-commercial, Higgs Audio carries a usage ceiling and mandatory attribution, Microsoft's VibeVoice stamps an audible “AI-generated” disclaimer into every file. Kokoro-82M passed everything: Apache-2.0, 82M parameters, a curated roster of preset narrators including three deep American men, and no watermark of any kind.

EngineSpeed on the M4Whole siteVoice sourceWatermark
Chatterbox2–8x SLOWER than realtime15–20 hclones any referenceyes (Perth)
Kokoro-82M6x FASTER than realtime~45 min54 curated presetsnone
What the speed number means

The full 43-piece corpus in a Kokoro voice is about 45 minutes of machine time, against 15–20 hours for Chatterbox and ~$5 for a cloud engine. And Kokoro carries no watermark, which removes the one standing reservation about publishing local audio. If one of its three voices below sounds right, the whole economics question collapses: free, fast, clean, and the voice can never be deprecated.

What Kokoro gives up: it cannot clone, so “exactly Brian” is not on its menu — only its own narrators. And its delivery is steadier than expressive; for essays that is arguably right.

03 · EvidenceThe checks

TakeTranscribesNote
Brian cloned · straightword perfect
Brian cloned · gravitasword perfect
Brian cloned · bass shelf“Thermust”the EQ'd reference bends the opening
Kokoro Onyxword perfect
Kokoro Fenrirword perfect
Kokoro Michaelword perfect

One honesty note on the cloning row: Brian is an ElevenLabs catalogue voice, and cloning a vendor's voice into another engine to sidestep their pricing is the kind of move worth doing with eyes open. For an experiment page it is a fair test of the local ceiling; as the shipped site voice, a Kokoro preset or a licensed render is the cleaner path.

04 · NextThe pick that ends the search

If a Kokoro voice wins, the full site renders in under an hour, watermark-free, and the reader ships. If a Brian clone wins, the corpus costs 15–20 hours of machine time and keeps the watermark question. If only Brian himself wins, the site voice is a paid render — ~$10–20 pay-as-you-go, one time, files yours forever. Name the row.