CosyVoice
the best permissive multilingual open model
the strongest permissively-licensed multilingual open model: apache-2.0, streaming, zero-shot cloning, and genuinely good across asian languages and english. documentation is research-grade and it's heavier to productionise than kokoro.
cosyvoice, from alibaba's funaudiollm, is multilingual (cjk plus english and more) with streaming, zero-shot voice cloning, and cross-lingual and instruction-following synthesis in its 2/3 line — and crucially it's apache-2.0, so it's commercial-safe. for multilingual self-hosting, especially where asian languages matter, it's the best open option.
the trade-offs are practical: the documentation is research-grade and it's heavier to stand up than a tiny model like kokoro. alibaba also offers hosted variants via dashscope if you'd rather not run it yourself (verify that pricing separately).
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- ai text-to-speech models
- pricing
- free — Apache-2.0 (weights included)
- website
- github.com
the strongest permissively-licensed multilingual open model: apache-2.0, streaming, zero-shot cloning, and genuinely good across asian languages and english. documentation is research-grade and it's heavier to productionise than kokoro.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.