Chatterbox
the best truly commercial-licensed open voice-clone model
the open model to reach for when you need voice cloning and a licence you can ship: mit code and weights, actively maintained, with built-in watermarking. self-host gpu ops are on you, and quality claims are self-reported.
chatterbox, from resemble ai, does zero-shot voice cloning from about ten seconds of audio, with exaggeration/emotion control and built-in perth watermarking, and — the differentiator versus f5 and fish — both its code and weights are mit-licensed, so it's genuinely commercial-safe. it's very actively maintained and available self-hosted or via hosts like fal.ai and wiro. a multilingual variant exists.
resemble self-reports that listeners prefer it over elevenlabs, which is a vendor a/b test, not an independent result — check tts arena v2 instead. self-hosting means gpu ops. as the open cloning model with a clean licence, it's the standout.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- ai text-to-speech models
- pricing
- free — MIT (weights included)
- website
- github.com
the open model to reach for when you need voice cloning and a licence you can ship: mit code and weights, actively maintained, with built-in watermarking. self-host gpu ops are on you, and quality claims are self-reported.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.