verifier.org

best 16 ai text-to-speech models

ai voice models and apis ranked for developers — on price you can read, a licence you can actually ship on, and whether they stream fast enough for a voice agent.

last reviewed 23 jul 2026 · 16 tools tested ·list curated by Onur Ozcanxin

the short version
best overallElevenLabsteams that want top quality and the best sdks without running infrastructure93/100runner-upCartesia Sonicreal-time voice agents that need fast first-audio and per-usage pricing89/100best free optionKokorohigh-volume, non-cloning narration you want to self-host cheaply82/100

there are two markets wearing one label here. proprietary apis — elevenlabs, openai, google, cartesia, deepgram, azure, polly and the rest — sell you latency, uptime, sdks and, crucially, a commercial licence you never have to think about. open-weight models — kokoro, chatterbox, cosyvoice, fish/openaudio, orpheus and friends — sell you zero marginal cost and full control, in exchange for running gpus and reading licences very carefully.

the single biggest trap for a developer is the licence, not the model. several of the best-sounding open models ship non-commercial weights even though their code is mit or apache: f5-tts weights are cc-by-nc, fish/openaudio s1 weights are cc-by-nc-sa, and the old coqui xtts-v2 weights are under a non-commercial licence. a repo's github badge describes the code, not the weights, and routinely lies about what you're allowed to ship. any 'we'll self-host free open tts' plan has to clear the weights licence first — so we checked each one and label the commercial-safe ones plainly.

second reality: for voice agents — the fastest-growing use case — the ranking isn't 'who sounds best', it's 'who streams first-audio fastest and bills per minute of conversation'. that pushes cartesia, deepgram and rime up for real-time work and pushes studio-grade narration tools down. pick for your use case, not for a leaderboard.

on quality: we deliberately don't print elo scores, because the honest sources for them — tts arena v2 on hugging face and the artificial analysis speech arena — are preference-vote leaderboards that move weekly. where a vendor claims it 'beats' a rival, we mark it as self-reported, not independent.

advertisement
  1. 1

    ElevenLabs

    the quality and ecosystem leader, agents to narration

    93/100

    verdictthe default when you want top-tier voice without ops: flash covers agents, v2/v3 cover expressive narration, and the sdks and voice library are the best documented in the category. priciest per character of the majors.

    best for
    teams that want top quality and the best sdks without running infrastructure
    price
    $0.05 per 1K chars (Flash/Turbo v2.5)
    pricing note
    flash/turbo v2.5 $0.05/1k chars; multilingual v2 and v3 $0.10/1k chars; free plan included
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    instant + professional
    streaming
    yes (websocket)
    languages
    30+

    elevenlabs spans the whole range — flash and turbo v2.5 for low-latency agent use, multilingual v2 and the expressive v3 for narration — across 30+ languages, with instant and professional voice cloning, websocket streaming, and a conversational-ai stack for real-time agents. the sdks, voice library and documentation are best-in-class, which is a real reason it's the safe default.

    the trade-offs: it's the priciest per character among the majors, and flash's ~75ms latency is a vendor claim with no independent benchmark. if you want quality with zero infrastructure, it's the tool; if you're cost-sensitive at high volume, weigh the open models below.

    pros
    • +top-tier quality across agent and narration models
    • +best-documented sdks and largest voice library
    • +instant and professional voice cloning, streaming
    cons
    • priciest per character of the majors
    • latency figures are vendor claims
    • closed — no self-host option
  2. 2

    Cartesia Sonic

    the voice-agent default: streaming-first, low latency

    89/100

    verdictthe go-to for voice agents: streaming-first by design with a vendor-claimed sub-100ms latency and clean credit pricing. quality is very good though not always top of the arena — check tts arena v2 for current standing.

    best for
    real-time voice agents that need fast first-audio and per-usage pricing
    price
    free tier; from $5/mo (Pro)
    pricing note
    billed 1 credit per character; free 20k/mo, pro $5/mo 100k, startup $49/mo 1.25m, scale $299/mo 8m
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    instant (pro tier)
    streaming
    yes (real-time focus)
    languages
    multilingual

    cartesia's sonic line is built for real-time: streaming-first architecture, instant voice cloning, and low latency by design, billed at one credit per character with a genuine free tier and plans scaling from $5/mo. for building a voice agent where time-to-first-audio is the metric that matters, it's the most natural pick in the category.

    the honest caveats: the credit accounting is more opaque than a flat per-character rate, its sub-100ms claim is unbenchmarked, and while quality is very good it isn't always the arena leader. for agents specifically, that's usually the right trade.

    pros
    • +streaming-first, built for real-time agents
    • +clean per-credit pricing with a free tier
    • +instant voice cloning
    cons
    • credit accounting less transparent than flat $/char
    • sub-100ms latency is a vendor claim
    • not always the top of the quality arenas
  3. 3

    OpenAI TTS

    cheap, steerable, and one client if you're already on openai

    86/100

    verdictthe best ergonomics if you're already on the openai sdk: one client, same auth, and gpt-4o-mini-tts is steerable by natural-language instructions. no voice cloning, and token pricing makes forecasting fuzzy.

    best for
    teams already building on openai who want cheap, steerable narration
    price
    $15 per 1M chars (tts-1)
    pricing note
    tts-1 $15/1m chars, tts-1-hd $30/1m; gpt-4o-mini-tts is token-priced (~$0.015/min per openai)
    free tier
    no
    type
    proprietary api
    license
    commercial api
    voice cloning
    no
    streaming
    yes
    languages
    multilingual (english-centric)

    openai offers tts-1 and tts-1-hd at flat per-character rates and gpt-4o-mini-tts, a token-priced model you can steer with natural-language 'instructions' for tone and emotion — something tts-1 can't do. it ships around eleven voices with streaming, and the killer feature is integration: if you're already on the openai sdk it's one client and one auth.

    the limits are clear: no voice cloning, fewer voices than the specialists, english-centric though multilingual works, and token-priced audio makes cost forecasting imprecise. for existing openai builders who want cheap, steerable speech, it's an easy add.

    pros
    • +excellent if already on the openai sdk
    • +gpt-4o-mini-tts steerable via natural-language instructions
    • +cheap flat per-character options
    cons
    • no voice cloning
    • fewer voices than the specialists
    • token pricing makes cost forecasting fuzzy
    advertisement
  4. 4

    Deepgram Aura-2

    purpose-built for english voice agents

    84/100

    verdictbuilt for conversational agents: simple per-character pricing, low-latency streaming, generous free credit, and it pairs with deepgram stt for a full agent loop. narrower on languages, not a cloning tool.

    best for
    english voice agents and ivr where one vendor does stt and tts
    price
    $0.030 per 1K chars (Aura-2)
    pricing note
    aura-1 $0.015/1k chars, aura-2 $0.030/1k; $200 free credit for new accounts
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    limited
    streaming
    yes (low latency)
    languages
    english-focused

    deepgram's aura-2 is a tts model aimed squarely at english voice agents — a large voice set, low-latency streaming, and clean rest/websocket apis, with $200 in free credit for new accounts. its real advantage is pairing with deepgram's speech-to-text so your whole agent loop (listen and speak) comes from one vendor with one integration.

    it's english-focused with narrower language coverage than google or azure, its sub-250ms latency claim is unbenchmarked, and it isn't a cloning or narration tool. for english ivr and voice agents where stt+tts unity matters, it's a strong, simple choice.

    pros
    • +purpose-built for english voice agents
    • +simple per-character pricing, generous free credit
    • +pairs with deepgram stt for a full agent loop
    cons
    • language coverage narrower than google/azure
    • not a cloning or narration tool
    • latency figure is a vendor claim
  5. 5

    Google Gemini TTS + Cloud TTS

    the widest catalogue and the cheapest floor

    83/100

    verdictunmatched breadth — seven voice tiers plus token-priced gemini tts — and the cheapest entry point at $4 per million characters, with a generous free tier. the multi-tier catalogue is confusing and some gemini models are preview.

    best for
    gcp-native teams wanting breadth and the lowest price floor
    price
    $4 per 1M chars (Standard/WaveNet)
    pricing note
    cloud tts standard/wavenet $4/1m, neural2 $16/1m, chirp3-hd $30/1m, studio $160/1m; gemini tts is token-priced; generous free tier
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    gated custom voice
    streaming
    yes
    languages
    40+

    google spans a huge catalogue: cloud text-to-speech from $4/million characters (standard/wavenet) up through neural2, chirp 3 hd and studio, plus the newer prompt-controllable, multi-speaker gemini tts (token-priced), across 40+ languages with ssml and a large free tier. for breadth and a low price floor, nothing here beats it.

    the friction is the sprawl — the multi-tier catalogue is genuinely confusing to price — and some gemini tts models carry preview labels, so confirm general availability before you commit. cloning is limited (custom voice is a separate gated program). for gcp-native shops it's an obvious pick.

    pros
    • +widest voice catalogue and 40+ languages
    • +cheapest floor at $4/1m characters
    • +generous free tier; gemini tts adds prompt control
    cons
    • multi-tier catalogue is confusing to price
    • some gemini tts models are preview, not ga
    • cloning limited to a separate gated program
  6. 6

    Kokoro

    the self-hoster's default: tiny, cpu-capable, truly permissive

    82/100

    verdictthe best cost/quality self-host for high-volume narration: 82m params, runs on cpu, and — rare in open tts — both code and weights are apache-2.0, so you can actually ship it. no cloning, and commit cadence has slowed.

    best for
    high-volume, non-cloning narration you want to self-host cheaply
    price
    free — Apache-2.0 (weights included)
    pricing note
    both code and the kokoro-82m weights are apache-2.0 — genuinely commercial-safe; hosted on fal.ai at $0.02/1k chars
    free tier
    yes
    type
    open-weight
    license
    Apache-2.0 (code + weights)
    voice cloning
    no (fixed voices)
    streaming
    yes
    languages
    english + some multilingual

    kokoro is an 82-million-parameter model small enough to run on cpu, with several english and some multilingual voices, and the thing that sets it apart from the higher-quality open models is the licence: both the code and the kokoro-82m weights are apache-2.0, so it's genuinely commercial-safe. self-host it for effectively nothing, or use it on fal.ai at $0.02 per 1k characters.

    it uses fixed voice packs, so there's no zero-shot cloning, and its expressiveness sits below the big proprietary models. its repo commit cadence has slowed (last substantive push in 2025), worth watching for staleness. for high-volume, non-cloning narration where cost matters most, it's the pick.

    pros
    • +tiny (82m) and cpu-capable
    • +code and weights both apache-2.0 — commercial-safe
    • +near-zero cost self-hosted
    cons
    • no zero-shot voice cloning (fixed voices)
    • expressiveness below the proprietary leaders
    • commit cadence has slowed
  7. 7

    Azure Neural TTS

    enterprise breadth, richest ssml, the safe corporate pick

    80/100

    verdictthe enterprise-safe choice: 140+ locales, the richest ssml in the category, custom neural voice and regional/compliance coverage. setup is heavier than the newcomers, and custom voice is gated.

    best for
    enterprises needing compliance, regions, and deep ssml control
    price
    $16 per 1M chars (Neural)
    pricing note
    neural $16/1m chars, neural hd $22/1m (reduced march 2026); commitment tiers lower; 500k chars/mo free
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    gated custom voice
    streaming
    yes
    languages
    140+ locales

    azure ai speech offers neural tts across 140+ languages and locales with the richest ssml controls (styles, roles, prosody), streaming, and a gated custom neural voice program, at $16/million characters (neural hd $22, reduced in march 2026) with commitment tiers lower and a 500k/month free tier. for enterprises that need compliance, data regions and fine ssml control, it's the safe corporate pick.

    the trade-off is weight: setup is heavier than the newer api-first tools, and custom voice requires approval. if you're already on azure and need governance, it fits naturally.

    pros
    • +140+ languages/locales, richest ssml
    • +enterprise compliance and regional coverage
    • +custom neural voice (gated) and streaming
    cons
    • heavier setup than newer api-first tools
    • custom voice is gated/approval-based
    • closed — no self-host
  8. 8

    Rime

    voice-agent specialist for telephony, with on-prem

    78/100

    verdictbuilt for high-call-volume agents: per-model per-character pricing, pronunciation control, and on-prem/vpc deployment most rivals don't offer. smaller brand and a narrower voice library than elevenlabs.

    best for
    contact-center and telephony voice agents needing pronunciation reliability
    price
    from $0.05 per 1K chars
    pricing note
    per-model rates reported around mist $0.03 / arcana $0.04 / coda $0.05 per 1k chars; 3,000 free minutes; enterprise on-prem/vpc
    free tier
    yes
    type
    proprietary api
    license
    commercial api (on-prem available)
    voice cloning
    limited
    streaming
    yes (low latency)
    languages
    10+ (arcana)

    rime targets telephony and voice agents with mist v2 for high volume and the expressive arcana flagship, offering pronunciation control, low latency, 3,000 free minutes to start, and — unusually — on-prem and vpc deployment for teams that can't send audio to a public api. for contact-center use where getting names and numbers right matters, that reliability focus is the draw.

    it's a smaller brand with a narrower voice library than the leaders, and its on-prem latency figures are vendor claims. for high-volume agent workloads needing pronunciation reliability and deployment control, it earns its spot.

    pros
    • +built for telephony/agents with pronunciation control
    • +on-prem and vpc deployment options
    • +per-model per-character pricing, free minutes
    cons
    • smaller brand, narrower voice library
    • latency figures are vendor claims
    • less suited to studio narration
  9. 9

    Chatterbox

    the best truly commercial-licensed open voice-clone model

    77/100

    verdictthe open model to reach for when you need voice cloning and a licence you can ship: mit code and weights, actively maintained, with built-in watermarking. self-host gpu ops are on you, and quality claims are self-reported.

    best for
    teams that need open voice cloning with a clean commercial licence
    price
    free — MIT (weights included)
    pricing note
    code and weights both mit — commercial-safe; also hosted on fal.ai and wiro
    free tier
    yes
    type
    open-weight
    license
    MIT (code + weights)
    voice cloning
    zero-shot (~10s)
    streaming
    yes
    languages
    english + multilingual variant

    chatterbox, from resemble ai, does zero-shot voice cloning from about ten seconds of audio, with exaggeration/emotion control and built-in perth watermarking, and — the differentiator versus f5 and fish — both its code and weights are mit-licensed, so it's genuinely commercial-safe. it's very actively maintained and available self-hosted or via hosts like fal.ai and wiro. a multilingual variant exists.

    resemble self-reports that listeners prefer it over elevenlabs, which is a vendor a/b test, not an independent result — check tts arena v2 instead. self-hosting means gpu ops. as the open cloning model with a clean licence, it's the standout.

    pros
    • +zero-shot cloning with mit code and weights
    • +very actively maintained
    • +built-in watermarking; multilingual variant
    cons
    • self-host gpu ops required
    • quality claims are self-reported, not independent
    • cloning raises consent/abuse considerations
  10. 10

    Amazon Polly

    the cheap, stable incumbent — boring in the best way

    75/100

    verdictthe dependable aws incumbent: 30+ languages, ssml, speech marks and a generative expressive tier, at low prices with a real free tier. top-line expressiveness trails elevenlabs and hume; no cloning.

    best for
    aws-native apps that want cheap, rock-stable tts
    price
    $4 per 1M chars (Standard)
    pricing note
    standard $4/1m, neural $16/1m, generative $30/1m, long-form $100/1m; 12-month free tier
    free tier
    yes
    type
    proprietary api
    license
    commercial api
    voice cloning
    no
    streaming
    yes
    languages
    30+

    polly is aws's steady tts: 30+ languages, ssml, speech marks, and a newer generative expressive tier, priced from $4/million characters (neural $16, generative $30) with a 12-month free tier and up to $200 in new-account credit. if you live in aws and want cheap, reliable speech that just works, it's the obvious pick.

    its top-line expressiveness trails the specialist leaders and there's no zero-shot cloning. that's the trade for stability and price — for many production apps, the right one.

    pros
    • +cheap and extremely stable
    • +30+ languages, ssml, speech marks
    • +generative expressive tier; real free tier
    cons
    • expressiveness trails elevenlabs/hume
    • no zero-shot cloning
    • best value only inside aws
  11. 11

    Hume Octave

    emotion-steerable tts you prompt in words

    73/100

    verdictthe model for expressive delivery: octave takes natural-language emotional prompts instead of ssml, so you describe the delivery in words. pricing is reported inconsistently across sources and the ecosystem is smaller.

    best for
    products where emotional delivery is the actual feature
    price
    from $3/mo (Starter)
    pricing note
    starter $3/mo (30k chars); business overage reported ~$0.05/1k chars; verify tier on hume.ai/pricing
    free tier
    no
    type
    proprietary api
    license
    commercial api
    voice cloning
    voice design
    streaming
    yes
    languages
    multilingual

    hume's octave was the first tts to take natural-language emotional prompts rather than ssml — you tell it how a line should be delivered, in words — with expressive, steerable output, voice design and multilingual support. for characters, companions and expressive narration where emotional delivery is the point, it does something the others don't.

    the caveats: pricing tiers are reported inconsistently across third parties (confirm on hume's own page), and its ecosystem is smaller than the majors. if emotional control is your feature, it's worth the look.

    pros
    • +natural-language emotional prompting (no ssml)
    • +expressive, steerable delivery and voice design
    • +multilingual
    cons
    • pricing reported inconsistently — verify live
    • smaller ecosystem than the majors
    • closed — no self-host
  12. 12

    CosyVoice

    the best permissive multilingual open model

    72/100

    verdictthe strongest permissively-licensed multilingual open model: apache-2.0, streaming, zero-shot cloning, and genuinely good across asian languages and english. documentation is research-grade and it's heavier to productionise than kokoro.

    best for
    multilingual self-hosting, especially cjk plus english, with cloning
    price
    free — Apache-2.0 (weights included)
    pricing note
    apache-2.0, commercial-safe; hosted variants via alibaba dashscope (verify separately)
    free tier
    yes
    type
    open-weight
    license
    Apache-2.0
    voice cloning
    zero-shot
    streaming
    yes
    languages
    multilingual (cjk + english)

    cosyvoice, from alibaba's funaudiollm, is multilingual (cjk plus english and more) with streaming, zero-shot voice cloning, and cross-lingual and instruction-following synthesis in its 2/3 line — and crucially it's apache-2.0, so it's commercial-safe. for multilingual self-hosting, especially where asian languages matter, it's the best open option.

    the trade-offs are practical: the documentation is research-grade and it's heavier to stand up than a tiny model like kokoro. alibaba also offers hosted variants via dashscope if you'd rather not run it yourself (verify that pricing separately).

    pros
    • +best permissively-licensed multilingual open model
    • +apache-2.0 — commercial-safe
    • +streaming and zero-shot cloning; strong on cjk
    cons
    • research-grade documentation
    • heavier to productionise than kokoro
    • self-host gpu ops
  13. 13

    Fish Speech / OpenAudio S1

    top-tier open quality — but the weights are non-commercial

    70/100

    verdicttop-tier open quality with zero-shot cloning and low latency — but the licence is the whole story: the s1 weights are non-commercial, so production means the paid hosted api, not the open weights.

    best for
    research and demos self-hosted; production via the paid hosted api
    price
    weights non-commercial (CC-BY-NC-SA); paid API for commercial use
    pricing note
    code is apache-2.0 but openaudio s1 weights are cc-by-nc-sa-4.0 — you cannot ship the weights commercially; fish.audio sells a paid api that is the commercial path
    free tier
    yes
    type
    open-weight + api
    license
    code Apache-2.0; weights CC-BY-NC-SA
    voice cloning
    zero-shot
    streaming
    yes
    languages
    multilingual

    fish speech / openaudio s1 delivers some of the best open-model quality here — zero-shot cloning, multilingual, low latency — and it's a common reason developers reach for open tts. but the trap is the licence: the code is apache-2.0 while the openaudio s1 weights are cc-by-nc-sa-4.0, non-commercial and share-alike, so you cannot ship the open weights in a commercial product.

    the commercial path is fish audio's paid api (per-character/credit on fish.audio — verify the current rate), not the downloadable weights. great for research and demos self-hosted; for production, budget for the api. this is the single place developers most often get burned in this category.

    pros
    • +top-tier open-model quality
    • +zero-shot cloning, multilingual, low latency
    • +paid hosted api available for commercial use
    cons
    • s1 weights are cc-by-nc-sa — non-commercial
    • commercial use requires the paid api, not the weights
    • self-host gpu ops for the free path
  14. 14

    Orpheus

    llm-style emotive tts with streaming, apache-licensed

    68/100

    verdicta llama-based, emotive open model with streaming and emotion tags, apache-2.0 so you can ship it — english-centric, and you'll need a real gpu for real-time.

    best for
    developers who want an expressive, apache-licensed model for agents
    price
    free — Apache-2.0 (weights included)
    pricing note
    apache-2.0, commercial-safe; hosted on fal.ai at $0.05/1k chars
    free tier
    yes
    type
    open-weight
    license
    Apache-2.0
    voice cloning
    limited
    streaming
    yes (~gpu required)
    languages
    english

    orpheus, from canopy labs, is a llama-3.2-3b-based tts with eight voices, emotive tags (laugh, sigh and the like), and low-latency streaming, all under apache-2.0 so it's commercial-safe. self-host it or use it on fal.ai at $0.05 per 1k characters. for an expressive, permissively-licensed open model aimed at agents, it's a good pick.

    it's english-centric and, being a 3b model, needs a real gpu for real-time use, so its ~200ms streaming figure assumes proper hardware and is a vendor claim. within those bounds it's one of the more expressive open options.

    pros
    • +expressive, emotive llm-style tts with streaming
    • +apache-2.0 — commercial-safe
    • +hosted option on fal.ai
    cons
    • english-centric
    • needs a real gpu for real-time
    • latency figure is a vendor claim
  15. 15

    Resemble AI

    per-second cloning with a deepfake-detection story

    66/100

    verdicthigh-quality cloning and real-time synthesis billed per second, paired with a deepfake-detection product line — and it's the org behind open-source chatterbox. per-second billing is awkward to compare, with clone add-on fees.

    best for
    products needing voice cloning plus an authenticity/security story
    price
    $0.0005 per second (~$1.80/hr of audio)
    pricing note
    tts $0.0005/sec, voice-agent path $0.001/sec; voice clones as monthly add-ons
    free tier
    no
    type
    proprietary api
    license
    commercial api
    voice cloning
    yes (add-on fees)
    streaming
    yes (real-time)
    languages
    multilingual

    resemble ai offers high-quality voice cloning, real-time synthesis and emotion control, billed unusually per second (about $1.80 per hour of audio; $0.001/sec on the voice-agent path), with a distinctive companion product line for deepfake detection and watermarking (resemble detect). it's also the company that open-sourced chatterbox, which speaks to its cloning expertise.

    the friction is commercial: per-second billing is hard to compare against per-character rivals, and voice clones are monthly add-on fees. for a product that needs cloning plus a security/authenticity narrative, it's a coherent choice.

    pros
    • +high-quality cloning and real-time synthesis
    • +deepfake-detection and watermarking product line
    • +the org behind open-source chatterbox
    cons
    • per-second billing awkward to compare
    • voice clones cost monthly add-on fees
    • no clear free tier
  16. 16

    VoxCPM

    the newest high-star expressive open model

    64/100

    verdicta young, tokenizer-free, context-aware open model with zero-shot cloning and an apache-2.0 licence you can ship — high stars, but fewer production references than kokoro or cosyvoice, so check the arenas rather than take claims on faith.

    best for
    developers wanting the newest permissive, expressive open model
    price
    free — Apache-2.0 (weights included)
    pricing note
    apache-2.0, commercial-safe; listed on wiro at $0.0001/char with a realtime voxcpm2 variant
    free tier
    yes
    type
    open-weight
    license
    Apache-2.0
    voice cloning
    zero-shot (context-aware)
    streaming
    yes (voxcpm2)
    languages
    multilingual

    voxcpm, from openbmb, is a compact (~0.5b) tokenizer-free model with context-aware zero-shot cloning and expressive output, apache-2.0 and therefore commercial-safe, with a realtime voxcpm2 variant and a hosted listing on wiro at $0.0001/char. it's the newest of the high-star open expressive models and an attractive option for developers who want something current and permissive.

    being young, it has fewer production references than kokoro or cosyvoice, and its quality standing should be checked on tts arena v2 rather than asserted. as a fresh, permissive, expressive open model, it rounds out the list.

    pros
    • +tokenizer-free, context-aware zero-shot cloning
    • +apache-2.0 — commercial-safe
    • +realtime variant; hosted option
    cons
    • young project, fewer production references
    • quality standing should be checked on the arenas
    • self-host gpu ops

how this ranking was made

prices come from each vendor's own pricing page, with the unit stated exactly as they state it — per 1k characters, per million, per second, or per token — because these are not comparable as printed. token-priced models genuinely can't be reduced to a clean per-character number, and we don't pretend otherwise.

for open-weight models, stars and last-commit dates were read from the github rest api on the review date, and the licence was verified by reading the actual file. we flag every non-commercial or gated weights licence, because that is where developers get burned.

latency figures are vendor marketing unless attributed to a named benchmark, and we found no independent reproducible latency benchmark across this whole field — so any millisecond number here is labelled a vendor claim and treated as directional.

quality is cited to a named arena, never asserted. vendor 'preferred over x%' stats are labelled self-reported a/b, not independent.

dead and renamed projects are dated: coqui shut down in january 2024 (and its popular xtts-v2 weights are non-commercial), and metavoice's team joined elevenlabs — both are flagged rather than ranked.

our general methodology and disclosures →
was this useful?