verifier.org

best 9 speech-to-text apis

every price converted to the same unit — dollars per hour of audio — plus the two things that change it: whether diarization costs extra, and whether the cheap rate assumes they can train on your recordings.

last reviewed 29 jul 2026 · 9 tools tested ·list curated by Onur Ozcanxin

the short version
best overallGladiaanyone processing audio they don't own — customer calls, interviews, medical or legal recordings.88/100runner-upAssemblyAIhigh-volume batch transcription where cost per hour is the deciding number.86/100best free optionElevenLabs Scribeaccuracy-critical transcription with lots of speakers — provided you set up billing correctly.85/100

vendors quote transcription in per-minute, per-second, per-character and per-api-call units, which makes the market look more differentiated than it is. converted to dollars per hour of audio, the managed services run from rev ai's $0.10 to gladia's $0.61 on pay-as-you-go — and two of the biggest names publish no readable figure at all.

the finding that reorders this ranking: deepgram's advertised rate assumes you let it train on your audio. standard pricing includes enrolment in its model improvement partnership programme, and opting out is a parameter you pass — which forfeits a discount worth roughly half the price. the headline number and the private-data number are not the same number, and only one of them is on the pricing page.

gladia goes the other way and makes no-training the contractual default, verifiable in its data processing agreement, while including diarization free at every tier. everyone else charges extra for speaker separation or leaves the training question to a settings toggle you have to find. assemblyai adds a subtler one: streaming is billed by how long the websocket stays open, not by how much audio passes through it, so silence costs the same as speech.

two of the three hyperscalers can't be priced at all. google's speech-to-text pricing page renders its tables through javascript and produced no figures across three attempts; azure's shows every price as a placeholder pending region selection, and the numbers circulating through microsoft's own community channels contradict each other by a factor of two. for a buyer, that is itself the review.

advertisement
  1. 1

    Gladia

    the only vendor here that makes not training on your audio the contractual default, with diarization thrown in free.

    88/100

    verdictthe most expensive entry rate here and the only data policy you don't have to negotiate for — which for most regulated work settles it.

    best for
    anyone processing audio they don't own — customer calls, interviews, medical or legal recordings.
    price
    $0.61/hr
    pricing note
    starter pay-as-you-go: $0.61/hr async, $0.75/hr real-time; growth tier as low as $0.20 and $0.25 with a volume commitment
    free tier
    yes
    cost/hr batch
    $0.61
    cost/hr realtime
    $0.75
    diarization
    included free
    trains on your audio
    no, by default
    self-host
    enterprise only

    gladia states that it does not train on your audio, that this is the default rather than an upgrade, and that it is verifiable in the data processing agreement. every rival either trains by default with an opt-out you have to find, or declines to say. if the recordings belong to your customers rather than to you, that difference is the product.

    diarization is included at every tier at no extra charge, where deepgram charges $0.12 an hour for it and assemblyai up to $0.12 on top of streaming. 100+ languages with mid-sentence language switching, and a self-reported 9.6% word error rate on real english audio for its solaria-3 model.

    the price is the trade. $0.61 an hour pay-as-you-go is the highest entry rate in this ranking — six times rev ai's cheapest tier — and the $0.20 figure on the marketing page needs a committed volume contract. read the headline as the committed price, not yours.

    pros
    • +no training on your audio, contractually, by default
    • +diarization included free at every tier
    • +100+ languages with mid-sentence switching
    • +one-time €50 credit with no monthly reset
    cons
    • $0.61/hr pay-as-you-go is the highest entry rate here
    • the advertised $0.20/hr needs a volume commitment
    • on-premise listed but detailed only via sales
    • word error rate is self-reported with no named dataset
  2. 2

    AssemblyAI

    cheapest managed batch transcription, quoted in the right unit — with streaming billed by the clock, not the audio.

    86/100

    verdictthe best price-to-capability ratio here and the only vendor already quoting in dollars per hour — spoiled by a streaming meter that runs on silence.

    best for
    high-volume batch transcription where cost per hour is the deciding number.
    price
    $0.15/hr
    pricing note
    universal-2 at $0.15/hr batch, universal-3.5 pro at $0.21/hr; real-time $0.45/hr base, diarization +$0.02 to $0.12/hr depending on mode
    free tier
    yes
    cost/hr batch
    $0.15-0.21
    cost/hr realtime
    $0.45 base
    diarization
    +$0.02 to $0.12/hr
    trains on your audio
    yes, opt-out available
    self-host
    not offered

    $0.15 an hour on universal-2 and $0.21 on the newer universal-3.5 pro are the cheapest fully-managed batch rates among the pure-play vendors, and assemblyai is refreshingly the only one that publishes in dollars per hour rather than making you convert. 99+ languages on streaming.

    the streaming gotcha is worth a paragraph of anyone's attention: real-time is billed by how long the websocket stays open, not by how much audio you send. an hour-long connection carrying thirty minutes of speech bills a full hour. for a call-centre application with hold music and silence, the effective rate is well above the quoted $0.45.

    on data, assemblyai's terms grant it a licence to use customer data including for model training, with an opt-out available through account settings or support. that is the common position in this category — better than deepgram's price-linked version, worse than gladia's default.

    pros
    • +cheapest managed batch rate at $0.15/hr
    • +already quotes in dollars per hour
    • +99+ languages on streaming
    • +$50 free credit with no card required
    cons
    • streaming billed by socket-open time, not audio
    • trains on your data by default; opt-out is manual
    • diarization costs extra in every mode
    • no self-hosting option at all
  3. 3

    ElevenLabs Scribe

    top of the independent accuracy leaderboard at $0.22 an hour — or nine times that if you pay the wrong way.

    85/100

    verdictthe most accurate option here by an independent measure, at a genuinely low rate, sitting behind a billing choice that can cost you nine times more for identical output.

    best for
    accuracy-critical transcription with lots of speakers — provided you set up billing correctly.
    price
    $0.22/hr
    pricing note
    pay-as-you-go $0.22/hr; paying with subscription credits costs about $1.98/hr, because scribe consumes 330 credits per minute
    free tier
    yes
    cost/hr batch
    $0.22 payg
    cost/hr realtime
    not separately published
    diarization
    included, 32 speakers
    trains on your audio
    not published
    self-host
    not offered

    scribe v2 ranks first for word error rate on artificial analysis's aa-agenttalk dataset at 1.5% — an independent leaderboard result rather than an elevenlabs claim, which is rare enough in this category to be the headline. diarization handles up to 32 speakers with per-word and per-character timestamps, plus audio-event tagging for things like laughter and music.

    the billing trap is severe and easy to walk into. pay-as-you-go is $0.22 an hour. paying with the credits bundled into a subscription costs about $1.98 an hour, because scribe burns 330 credits per minute and those credit pools are sized for text-to-speech character volume rather than transcription minutes. same model, same output, nine times the cost, and the plan page won't tell you.

    no self-hosting, and elevenlabs' training policy for scribe audio specifically wasn't confirmable on its pages — a gap worth closing yourself if the recordings are sensitive.

    pros
    • +1.5% word error rate, first on an independent leaderboard
    • +$0.22/hr on pay-as-you-go
    • +diarization for up to 32 speakers
    • +audio-event tagging alongside transcription
    cons
    • subscription credits cost roughly 9x the payg rate
    • training policy for scribe audio not published clearly
    • no self-hosting option
    • free tier is about 30 minutes a month
    advertisement
  4. 4

    OpenAI Whisper API

    the only entry whose model you can legally run yourself for free — which makes it the yardstick for everything else here.

    84/100

    verdictnot the most accurate and not the cheapest, but the only one where the managed price and the self-hosted price are the same model — and it doesn't train on your audio.

    best for
    teams who want a managed option today and the ability to walk away to their own hardware later.
    price
    $0.18/hr
    pricing note
    gpt-4o-mini-transcribe at $0.18/hr; whisper-1 and gpt-4o-transcribe at $0.36/hr; the whisper weights themselves are mit and free to self-host
    free tier
    no
    cost/hr batch
    $0.18-0.36
    cost/hr realtime
    not published
    diarization
    none native
    trains on your audio
    no, by default
    self-host
    yes, mit weights

    whisper's weights are mit-licensed across every size from tiny to large-v3, so the model openai bills $0.36 an hour for is one you can run on your own gpu for the cost of electricity. no other vendor here offers that exit, and it means every managed price on this page should be read as a convenience premium over a known free alternative.

    openai's api data controls state that audio sent to the transcription endpoints is not used for training unless you opt in, with roughly 30-day retention for abuse monitoring and stricter terms available. that is the correct default and only gladia matches it.

    two real gaps. there is no native speaker diarization at all — you bolt on a separate tool — and there's no free tier or trial credit for the transcription endpoints. worth knowing too that whisper large-v3 is no longer state of the art: the hugging face open asr leaderboard now places cohere transcribe and ibm granite speech above it, though whisper remains the most widely deployed baseline.

    pros
    • +mit weights, free to self-host, same model as the api
    • +no training on your audio by default
    • +gpt-4o-mini-transcribe at $0.18/hr
    • +99 languages
    cons
    • no native speaker diarization
    • no free tier or trial credit
    • no separately published real-time rate
    • whisper large-v3 no longer leads open leaderboards
  5. 5

    Rev AI

    ten cents an hour — the cheapest number in this ranking, attached to the least complete pricing page.

    80/100

    verdictless than half assemblyai's rate and a third of whisper's, with real-time pricing and diarization costs that simply aren't published.

    best for
    large batch archives where price per hour dominates and English is the main language.
    price
    $0.10/hr
    pricing note
    reverb turbo $0.10/hr, reverb $0.20/hr, foreign language and whisper-based models $0.30/hr; human transcription is a separate product at about $119/hr
    free tier
    yes
    cost/hr batch
    $0.10-0.30
    cost/hr realtime
    not published
    diarization
    claimed, billing unclear
    trains on your audio
    not published
    self-host
    claimed, unverified

    $0.10 an hour on reverb turbo is the lowest published rate anywhere in this ranking, with the standard reverb model at $0.20 and whisper-based options at $0.30. free credits equivalent to five hours of reverb make it cheap to evaluate.

    rev describes reverb and reverb turbo as open-source models, which would give it a self-hosting path alongside whisper's — but we could not confirm the licence or repository within budget, so treat that as a promising lead rather than a fact. its marketing cites gains 'over competitors' rather than an absolute error rate against a named dataset, so there's no accuracy figure here we'd publish.

    the pricing page is built around batch. no separate real-time rate appears despite a documented streaming api, and while diarization is marketed as part of the model, no line item confirms whether it's free or billed. 57+ languages async but only 9+ on streaming, which is a sharp drop.

    pros
    • +$0.10/hr — cheapest published rate here
    • +free credits worth five hours of transcription
    • +claimed open-source model lineage
    • +57+ languages for async transcription
    cons
    • no published real-time or streaming rate
    • diarization billing unclear
    • streaming supports only 9+ languages
    • open-source claim unverified — licence not confirmed
  6. 6

    Deepgram

    the clearest pricing table here, for a price that assumes your recordings become their training data.

    76/100

    verdictexcellent, itemised, honest-looking pricing — with the most important term about it disclosed in an api parameter rather than on the pricing page.

    best for
    teams whose audio isn't sensitive and who want streaming at the lowest real-time rate here.
    price
    $0.462/hr
    pricing note
    nova-3 monolingual $0.462/hr batch and $0.288/hr streaming; multilingual $0.552 and $0.348; diarization +$0.12/hr; opting out of model training forfeits a discount worth roughly half
    free tier
    yes
    cost/hr batch
    $0.462
    cost/hr realtime
    $0.288
    diarization
    +$0.12/hr
    trains on your audio
    yes — price depends on it
    self-host
    sales-gated

    the streaming rate of $0.288 an hour is the lowest real-time figure in this ranking, the pricing table itemises add-ons cleanly, and $200 of free credit makes evaluation easy. on presentation alone deepgram looks like the most transparent vendor here.

    then you find the condition. standard rates assume enrolment in deepgram's model improvement partnership programme — your audio helps train their models. you can opt out by passing mip_opt_out=true, and doing so forfeits the roughly 50% discount that enrolment buys. so the real choice is $0.288 an hour with training or about double without, and only the first number appears in the comparison tables everyone publishes.

    diarization is a further $0.12 an hour, where gladia includes it. on-premise deployment exists but pricing is entirely sales-gated, and nova-3 multilingual covers around 30 languages, fewer than most rivals here.

    pros
    • +lowest published real-time rate at $0.288/hr
    • +clearly itemised add-on pricing
    • +$200 free credit with no card
    • +steady multilingual expansion through 2026
    cons
    • headline price assumes you let deepgram train on your audio
    • opting out roughly doubles the effective cost
    • diarization costs an extra $0.12/hr
    • around 30 languages, fewer than rivals
  7. 7

    Speechmatics

    built for regulated on-premise deployment, with a pricing page that won't separate batch from real-time.

    70/100

    verdicta serious enterprise product whose self-serve pricing tells you less than any pure-play rival's — which matters less if you were always going to talk to sales.

    best for
    regulated enterprises that need on-premise or on-device deployment and will go through sales anyway.
    price
    $0.129/hr
    pricing note
    pro tier from $0.129/hr billed to the second; a $0.24/hr real-time figure appears elsewhere on the site but could not be confirmed, and the docs pricing pages returned 404
    free tier
    yes
    cost/hr batch
    from $0.129
    cost/hr realtime
    ~$0.24, unconfirmed
    diarization
    likely included, unconfirmed
    trains on your audio
    toggle, default unknown
    self-host
    yes, sales-gated

    the deployment options are the reason to look: saas, on-premises, and on-device, aimed squarely at buyers who cannot send audio to a third party. $100 of free credit with no card is generous for evaluation, billing is per second, and volume discounts apply automatically above 500 hours a month.

    the pricing page states 'from $0.129/hr' on the pro tier and then stops. it doesn't separate batch from real-time — the only vendor here that doesn't — and a $0.24 real-time figure we found elsewhere on the site couldn't be confirmed against a fetchable page. the accuracy documentation returned 404.

    the training default is likewise unresolved. data logging is described as a toggle customers can change at any time, but whether it starts on or off wasn't confirmable, which for a vendor selling to regulated buyers is the one thing that should be stated plainly.

    pros
    • +saas, on-premises and on-device deployment
    • +$100 free credit, no card required
    • +per-second billing with automatic volume discounts
    • +56+ languages
    cons
    • pricing page doesn't separate batch from real-time
    • accuracy documentation returns 404
    • default training setting unconfirmed
    • on-premise licensing requires sales
  8. 8

    Azure AI Speech

    genuinely offline containers for air-gapped work, and a price even microsoft's own channels can't agree on.

    64/100

    verdictthe best offline story in the category attached to the worst price transparency — you cannot find out what it costs without a sales conversation or a calculator session.

    best for
    air-gapped and regulated deployments where the disconnected container is the requirement.
    price
    unverified
    pricing note
    the pricing calculator renders every figure as a placeholder pending region selection; figures circulating via microsoft's own community channels conflict — roughly $0.18/hr against $0.36/hr for standard batch
    free tier
    yes
    cost/hr batch
    not readable
    cost/hr realtime
    not readable
    diarization
    free on batch, paid on realtime
    trains on your audio
    no, per azure policy
    self-host
    yes, offline containers

    the disconnected containers are a real differentiator. azure will let you run speech recognition fully offline in an air-gapped environment, subject to microsoft approval and a commitment plan, which almost no pure-play vendor offers. connected containers need no special approval. for defence, health and government work that can settle the decision on its own.

    the pricing is unreadable. every figure on the calculator is a '$-' placeholder until you pick a region and currency, and across three attempts nothing static could be extracted. the numbers that do circulate through microsoft's own community and support channels disagree with each other by a factor of two — about $0.18 an hour against about $0.36 for standard batch — which suggests an unannounced change nobody has documented.

    what we could confirm: the free f0 tier gives five audio hours a month of real-time only, batch isn't supported on it, and batch diarization is included at no extra charge on standard and custom tiers while real-time diarization is billed as an enhanced add-on.

    pros
    • +fully offline disconnected containers for air-gapped use
    • +batch diarization included at no extra charge
    • +free tier of five audio hours a month
    • +azure's enterprise compliance and support
    cons
    • no static price figure obtainable at all
    • microsoft's own sources give conflicting rates
    • batch not supported on the free tier
    • real-time diarization is a paid add-on
  9. 9

    Google Cloud Speech-to-Text

    chirp 3 behind a pricing page that is a calculator, with no number on it anywhere.

    60/100

    verdictcapable models and a real on-premise product line, sold from the least readable pricing page of the nine — we cannot tell you what it costs.

    best for
    teams already committed to google cloud, who will price it inside the console.
    price
    unverified
    pricing note
    the pricing page renders entirely through client-side javascript; three attempts including the v2 docs redirect produced no extractable dollar figure
    free tier
    yes
    cost/hr batch
    not readable
    cost/hr realtime
    not readable
    diarization
    available, pricing unclear
    trains on your audio
    no, per gcp policy
    self-host
    separate on-prem product

    the chirp 2 and chirp 3 models are competitive and the gcp integration is the deepest here for anyone already on google's stack. there's also a distinct on-premise product line, which most pure-play vendors don't offer at all.

    we could not extract a single price. the pricing page is calculator-driven and returned no static figures across three attempts, including following the v2 documentation redirect. the widely-quoted 60 free minutes a month couldn't be confirmed on google's own page either, and google's docs point you to an internal model-comparison tool rather than publishing a headline accuracy figure.

    one confirmed improvement worth noting: billing granularity moved from 15-second to 1-second rounding, which materially helps anyone transcribing lots of short clips. diarization is an available request parameter, though whether it carries a surcharge wasn't confirmable.

    pros
    • +chirp 2 and chirp 3 models with broad language coverage
    • +billing now rounds to the second, not 15 seconds
    • +dedicated on-premise product line
    • +deepest integration for existing gcp users
    cons
    • no price figure extractable from its own page
    • free-tier allowance unconfirmed on google's pages
    • no published headline accuracy figure
    • diarization surcharge status unclear

how this ranking was made

every price is converted to usd per hour of audio, with the arithmetic shown in the entry, because no two vendors quote the same unit. where batch and real-time differ we give both, and where a headline rate requires a volume commitment we quote the pay-as-you-go rate instead and note the committed price separately.

accuracy claims are reported as what they are. a vendor's own word error rate is labelled as the vendor's, with the dataset named where one is given. we deliberately exclude vendor-published comparisons against rivals — several vendors publish tables of their competitors' error rates, and those are marketing rather than measurement.

where an independent leaderboard has evaluated a model we cite it separately. elevenlabs' scribe v2 ranks first for word error rate on artificial analysis's aa-agenttalk set at 1.5%, and the hugging face open asr leaderboard — which only scores open-weight models — is currently led by cohere transcribe and ibm granite speech rather than by whisper, which is no longer state of the art despite remaining the default baseline.

the training-data policy is treated as a pricing field rather than a privacy footnote, because at deepgram it literally is one. we report the default, whether opting out is possible, and what opting out costs.

whisper is the reference point for what managed transcription is worth: the same model that openai bills at $0.36 an hour is mit-licensed and free to run on your own hardware. every managed price on this page should be read against that.

our general methodology and disclosures →
was this useful?