← A Voice Reader for Simply Broken
v6 · chasing brianCan free get to Brian?
Brian won v5. This round chases him with two local engines: Chatterbox cloning Brian's own clip, and Kokoro — a second engine, researched and installed today — whose native American narrator presets sidestep the accent problem entirely. Brian himself sits at the top of the rack as the target.
“Brian is great. Let's see if we can get to it locally. The other voices you generated locally sound Indian accent which is not good for my needs.”
Chatterbox clones its REFERENCE. Every earlier local take was cloned from your Hebrew recording, so the clones carried a non-native English accent — that is what read as Indian. The fix is in the reference: clone from Brian's clean American clip, or use an engine whose voices are native American English to begin with. Both are on this rack.
01 · ListenThe target, then the chasers
02 · The second engineWhy Kokoro, and what it changes
The research pass ranked the 2026 open-model field against three hard requirements: commercial licence, Apple Silicon, top-tier English. Most of the acclaimed names fell at a gate — Fish Speech's weights are non-commercial, Higgs Audio carries a usage ceiling and mandatory attribution, Microsoft's VibeVoice stamps an audible “AI-generated” disclaimer into every file. Kokoro-82M passed everything: Apache-2.0, 82M parameters, a curated roster of preset narrators including three deep American men, and no watermark of any kind.
| Engine | Speed on the M4 | Whole site | Voice source | Watermark |
|---|---|---|---|---|
| Chatterbox | 2–8x SLOWER than realtime | 15–20 h | clones any reference | yes (Perth) |
| Kokoro-82M | 6x FASTER than realtime | ~45 min | 54 curated presets | none |
The full 43-piece corpus in a Kokoro voice is about 45 minutes of machine time, against 15–20 hours for Chatterbox and ~$5 for a cloud engine. And Kokoro carries no watermark, which removes the one standing reservation about publishing local audio. If one of its three voices below sounds right, the whole economics question collapses: free, fast, clean, and the voice can never be deprecated.
What Kokoro gives up: it cannot clone, so “exactly Brian” is not on its menu — only its own narrators. And its delivery is steadier than expressive; for essays that is arguably right.
03 · EvidenceThe checks
| Take | Transcribes | Note |
|---|---|---|
| Brian cloned · straight | word perfect | — |
| Brian cloned · gravitas | word perfect | — |
| Brian cloned · bass shelf | “Thermust” | the EQ'd reference bends the opening |
| Kokoro Onyx | word perfect | — |
| Kokoro Fenrir | word perfect | — |
| Kokoro Michael | word perfect | — |
One honesty note on the cloning row: Brian is an ElevenLabs catalogue voice, and cloning a vendor's voice into another engine to sidestep their pricing is the kind of move worth doing with eyes open. For an experiment page it is a fair test of the local ceiling; as the shipped site voice, a Kokoro preset or a licensed render is the cleaner path.
04 · NextThe pick that ends the search
If a Kokoro voice wins, the full site renders in under an hour, watermark-free, and the reader ships. If a Brian clone wins, the corpus costs 15–20 hours of machine time and keeps the watermark question. If only Brian himself wins, the site voice is a paid render — ~$10–20 pay-as-you-go, one time, files yours forever. Name the row.