every price converted to the same unit — dollars per hour of audio — plus the two things that change it: whether diarization costs extra, and whether the cheap rate assumes they can train on your recordings.
last reviewed 29 jul 2026 · 9 tools tested ·list curated by Onur Ozcanxin
vendors quote transcription in per-minute, per-second, per-character and per-api-call units, which makes the market look more differentiated than it is. converted to dollars per hour of audio, the managed services run from rev ai's $0.10 to gladia's $0.61 on pay-as-you-go — and two of the biggest names publish no readable figure at all.
the finding that reorders this ranking: deepgram's advertised rate assumes you let it train on your audio. standard pricing includes enrolment in its model improvement partnership programme, and opting out is a parameter you pass — which forfeits a discount worth roughly half the price. the headline number and the private-data number are not the same number, and only one of them is on the pricing page.
gladia goes the other way and makes no-training the contractual default, verifiable in its data processing agreement, while including diarization free at every tier. everyone else charges extra for speaker separation or leaves the training question to a settings toggle you have to find. assemblyai adds a subtler one: streaming is billed by how long the websocket stays open, not by how much audio passes through it, so silence costs the same as speech.
two of the three hyperscalers can't be priced at all. google's speech-to-text pricing page renders its tables through javascript and produced no figures across three attempts; azure's shows every price as a placeholder pending region selection, and the numbers circulating through microsoft's own community channels contradict each other by a factor of two. for a buyer, that is itself the review.
advertisement
1
Gladia
the only vendor here that makes not training on your audio the contractual default, with diarization thrown in free.
88/100
verdictthe most expensive entry rate here and the only data policy you don't have to negotiate for — which for most regulated work settles it.
best for
anyone processing audio they don't own — customer calls, interviews, medical or legal recordings.
price
$0.61/hr
pricing note
starter pay-as-you-go: $0.61/hr async, $0.75/hr real-time; growth tier as low as $0.20 and $0.25 with a volume commitment
free tier
yes
cost/hr batch
$0.61
cost/hr realtime
$0.75
diarization
included free
trains on your audio
no, by default
self-host
enterprise only
gladia states that it does not train on your audio, that this is the default rather than an upgrade, and that it is verifiable in the data processing agreement. every rival either trains by default with an opt-out you have to find, or declines to say. if the recordings belong to your customers rather than to you, that difference is the product.
diarization is included at every tier at no extra charge, where deepgram charges $0.12 an hour for it and assemblyai up to $0.12 on top of streaming. 100+ languages with mid-sentence language switching, and a self-reported 9.6% word error rate on real english audio for its solaria-3 model.
the price is the trade. $0.61 an hour pay-as-you-go is the highest entry rate in this ranking — six times rev ai's cheapest tier — and the $0.20 figure on the marketing page needs a committed volume contract. read the headline as the committed price, not yours.
pros
+no training on your audio, contractually, by default
+diarization included free at every tier
+100+ languages with mid-sentence switching
+one-time €50 credit with no monthly reset
cons
−$0.61/hr pay-as-you-go is the highest entry rate here
−the advertised $0.20/hr needs a volume commitment
−on-premise listed but detailed only via sales
−word error rate is self-reported with no named dataset
cheapest managed batch transcription, quoted in the right unit — with streaming billed by the clock, not the audio.
86/100
verdictthe best price-to-capability ratio here and the only vendor already quoting in dollars per hour — spoiled by a streaming meter that runs on silence.
best for
high-volume batch transcription where cost per hour is the deciding number.
price
$0.15/hr
pricing note
universal-2 at $0.15/hr batch, universal-3.5 pro at $0.21/hr; real-time $0.45/hr base, diarization +$0.02 to $0.12/hr depending on mode
free tier
yes
cost/hr batch
$0.15-0.21
cost/hr realtime
$0.45 base
diarization
+$0.02 to $0.12/hr
trains on your audio
yes, opt-out available
self-host
not offered
$0.15 an hour on universal-2 and $0.21 on the newer universal-3.5 pro are the cheapest fully-managed batch rates among the pure-play vendors, and assemblyai is refreshingly the only one that publishes in dollars per hour rather than making you convert. 99+ languages on streaming.
the streaming gotcha is worth a paragraph of anyone's attention: real-time is billed by how long the websocket stays open, not by how much audio you send. an hour-long connection carrying thirty minutes of speech bills a full hour. for a call-centre application with hold music and silence, the effective rate is well above the quoted $0.45.
on data, assemblyai's terms grant it a licence to use customer data including for model training, with an opt-out available through account settings or support. that is the common position in this category — better than deepgram's price-linked version, worse than gladia's default.
pros
+cheapest managed batch rate at $0.15/hr
+already quotes in dollars per hour
+99+ languages on streaming
+$50 free credit with no card required
cons
−streaming billed by socket-open time, not audio
−trains on your data by default; opt-out is manual
top of the independent accuracy leaderboard at $0.22 an hour — or nine times that if you pay the wrong way.
85/100
verdictthe most accurate option here by an independent measure, at a genuinely low rate, sitting behind a billing choice that can cost you nine times more for identical output.
best for
accuracy-critical transcription with lots of speakers — provided you set up billing correctly.
price
$0.22/hr
pricing note
pay-as-you-go $0.22/hr; paying with subscription credits costs about $1.98/hr, because scribe consumes 330 credits per minute
free tier
yes
cost/hr batch
$0.22 payg
cost/hr realtime
not separately published
diarization
included, 32 speakers
trains on your audio
not published
self-host
not offered
scribe v2 ranks first for word error rate on artificial analysis's aa-agenttalk dataset at 1.5% — an independent leaderboard result rather than an elevenlabs claim, which is rare enough in this category to be the headline. diarization handles up to 32 speakers with per-word and per-character timestamps, plus audio-event tagging for things like laughter and music.
the billing trap is severe and easy to walk into. pay-as-you-go is $0.22 an hour. paying with the credits bundled into a subscription costs about $1.98 an hour, because scribe burns 330 credits per minute and those credit pools are sized for text-to-speech character volume rather than transcription minutes. same model, same output, nine times the cost, and the plan page won't tell you.
no self-hosting, and elevenlabs' training policy for scribe audio specifically wasn't confirmable on its pages — a gap worth closing yourself if the recordings are sensitive.
pros
+1.5% word error rate, first on an independent leaderboard
+$0.22/hr on pay-as-you-go
+diarization for up to 32 speakers
+audio-event tagging alongside transcription
cons
−subscription credits cost roughly 9x the payg rate
−training policy for scribe audio not published clearly
the only entry whose model you can legally run yourself for free — which makes it the yardstick for everything else here.
84/100
verdictnot the most accurate and not the cheapest, but the only one where the managed price and the self-hosted price are the same model — and it doesn't train on your audio.
best for
teams who want a managed option today and the ability to walk away to their own hardware later.
price
$0.18/hr
pricing note
gpt-4o-mini-transcribe at $0.18/hr; whisper-1 and gpt-4o-transcribe at $0.36/hr; the whisper weights themselves are mit and free to self-host
free tier
no
cost/hr batch
$0.18-0.36
cost/hr realtime
not published
diarization
none native
trains on your audio
no, by default
self-host
yes, mit weights
whisper's weights are mit-licensed across every size from tiny to large-v3, so the model openai bills $0.36 an hour for is one you can run on your own gpu for the cost of electricity. no other vendor here offers that exit, and it means every managed price on this page should be read as a convenience premium over a known free alternative.
openai's api data controls state that audio sent to the transcription endpoints is not used for training unless you opt in, with roughly 30-day retention for abuse monitoring and stricter terms available. that is the correct default and only gladia matches it.
two real gaps. there is no native speaker diarization at all — you bolt on a separate tool — and there's no free tier or trial credit for the transcription endpoints. worth knowing too that whisper large-v3 is no longer state of the art: the hugging face open asr leaderboard now places cohere transcribe and ibm granite speech above it, though whisper remains the most widely deployed baseline.
pros
+mit weights, free to self-host, same model as the api
+no training on your audio by default
+gpt-4o-mini-transcribe at $0.18/hr
+99 languages
cons
−no native speaker diarization
−no free tier or trial credit
−no separately published real-time rate
−whisper large-v3 no longer leads open leaderboards
ten cents an hour — the cheapest number in this ranking, attached to the least complete pricing page.
80/100
verdictless than half assemblyai's rate and a third of whisper's, with real-time pricing and diarization costs that simply aren't published.
best for
large batch archives where price per hour dominates and English is the main language.
price
$0.10/hr
pricing note
reverb turbo $0.10/hr, reverb $0.20/hr, foreign language and whisper-based models $0.30/hr; human transcription is a separate product at about $119/hr
free tier
yes
cost/hr batch
$0.10-0.30
cost/hr realtime
not published
diarization
claimed, billing unclear
trains on your audio
not published
self-host
claimed, unverified
$0.10 an hour on reverb turbo is the lowest published rate anywhere in this ranking, with the standard reverb model at $0.20 and whisper-based options at $0.30. free credits equivalent to five hours of reverb make it cheap to evaluate.
rev describes reverb and reverb turbo as open-source models, which would give it a self-hosting path alongside whisper's — but we could not confirm the licence or repository within budget, so treat that as a promising lead rather than a fact. its marketing cites gains 'over competitors' rather than an absolute error rate against a named dataset, so there's no accuracy figure here we'd publish.
the pricing page is built around batch. no separate real-time rate appears despite a documented streaming api, and while diarization is marketed as part of the model, no line item confirms whether it's free or billed. 57+ languages async but only 9+ on streaming, which is a sharp drop.
pros
+$0.10/hr — cheapest published rate here
+free credits worth five hours of transcription
+claimed open-source model lineage
+57+ languages for async transcription
cons
−no published real-time or streaming rate
−diarization billing unclear
−streaming supports only 9+ languages
−open-source claim unverified — licence not confirmed
the clearest pricing table here, for a price that assumes your recordings become their training data.
76/100
verdictexcellent, itemised, honest-looking pricing — with the most important term about it disclosed in an api parameter rather than on the pricing page.
best for
teams whose audio isn't sensitive and who want streaming at the lowest real-time rate here.
price
$0.462/hr
pricing note
nova-3 monolingual $0.462/hr batch and $0.288/hr streaming; multilingual $0.552 and $0.348; diarization +$0.12/hr; opting out of model training forfeits a discount worth roughly half
free tier
yes
cost/hr batch
$0.462
cost/hr realtime
$0.288
diarization
+$0.12/hr
trains on your audio
yes — price depends on it
self-host
sales-gated
the streaming rate of $0.288 an hour is the lowest real-time figure in this ranking, the pricing table itemises add-ons cleanly, and $200 of free credit makes evaluation easy. on presentation alone deepgram looks like the most transparent vendor here.
then you find the condition. standard rates assume enrolment in deepgram's model improvement partnership programme — your audio helps train their models. you can opt out by passing mip_opt_out=true, and doing so forfeits the roughly 50% discount that enrolment buys. so the real choice is $0.288 an hour with training or about double without, and only the first number appears in the comparison tables everyone publishes.
diarization is a further $0.12 an hour, where gladia includes it. on-premise deployment exists but pricing is entirely sales-gated, and nova-3 multilingual covers around 30 languages, fewer than most rivals here.
pros
+lowest published real-time rate at $0.288/hr
+clearly itemised add-on pricing
+$200 free credit with no card
+steady multilingual expansion through 2026
cons
−headline price assumes you let deepgram train on your audio
built for regulated on-premise deployment, with a pricing page that won't separate batch from real-time.
70/100
verdicta serious enterprise product whose self-serve pricing tells you less than any pure-play rival's — which matters less if you were always going to talk to sales.
best for
regulated enterprises that need on-premise or on-device deployment and will go through sales anyway.
price
$0.129/hr
pricing note
pro tier from $0.129/hr billed to the second; a $0.24/hr real-time figure appears elsewhere on the site but could not be confirmed, and the docs pricing pages returned 404
free tier
yes
cost/hr batch
from $0.129
cost/hr realtime
~$0.24, unconfirmed
diarization
likely included, unconfirmed
trains on your audio
toggle, default unknown
self-host
yes, sales-gated
the deployment options are the reason to look: saas, on-premises, and on-device, aimed squarely at buyers who cannot send audio to a third party. $100 of free credit with no card is generous for evaluation, billing is per second, and volume discounts apply automatically above 500 hours a month.
the pricing page states 'from $0.129/hr' on the pro tier and then stops. it doesn't separate batch from real-time — the only vendor here that doesn't — and a $0.24 real-time figure we found elsewhere on the site couldn't be confirmed against a fetchable page. the accuracy documentation returned 404.
the training default is likewise unresolved. data logging is described as a toggle customers can change at any time, but whether it starts on or off wasn't confirmable, which for a vendor selling to regulated buyers is the one thing that should be stated plainly.
pros
+saas, on-premises and on-device deployment
+$100 free credit, no card required
+per-second billing with automatic volume discounts
+56+ languages
cons
−pricing page doesn't separate batch from real-time
genuinely offline containers for air-gapped work, and a price even microsoft's own channels can't agree on.
64/100
verdictthe best offline story in the category attached to the worst price transparency — you cannot find out what it costs without a sales conversation or a calculator session.
best for
air-gapped and regulated deployments where the disconnected container is the requirement.
price
unverified
pricing note
the pricing calculator renders every figure as a placeholder pending region selection; figures circulating via microsoft's own community channels conflict — roughly $0.18/hr against $0.36/hr for standard batch
free tier
yes
cost/hr batch
not readable
cost/hr realtime
not readable
diarization
free on batch, paid on realtime
trains on your audio
no, per azure policy
self-host
yes, offline containers
the disconnected containers are a real differentiator. azure will let you run speech recognition fully offline in an air-gapped environment, subject to microsoft approval and a commitment plan, which almost no pure-play vendor offers. connected containers need no special approval. for defence, health and government work that can settle the decision on its own.
the pricing is unreadable. every figure on the calculator is a '$-' placeholder until you pick a region and currency, and across three attempts nothing static could be extracted. the numbers that do circulate through microsoft's own community and support channels disagree with each other by a factor of two — about $0.18 an hour against about $0.36 for standard batch — which suggests an unannounced change nobody has documented.
what we could confirm: the free f0 tier gives five audio hours a month of real-time only, batch isn't supported on it, and batch diarization is included at no extra charge on standard and custom tiers while real-time diarization is billed as an enhanced add-on.
pros
+fully offline disconnected containers for air-gapped use
chirp 3 behind a pricing page that is a calculator, with no number on it anywhere.
60/100
verdictcapable models and a real on-premise product line, sold from the least readable pricing page of the nine — we cannot tell you what it costs.
best for
teams already committed to google cloud, who will price it inside the console.
price
unverified
pricing note
the pricing page renders entirely through client-side javascript; three attempts including the v2 docs redirect produced no extractable dollar figure
free tier
yes
cost/hr batch
not readable
cost/hr realtime
not readable
diarization
available, pricing unclear
trains on your audio
no, per gcp policy
self-host
separate on-prem product
the chirp 2 and chirp 3 models are competitive and the gcp integration is the deepest here for anyone already on google's stack. there's also a distinct on-premise product line, which most pure-play vendors don't offer at all.
we could not extract a single price. the pricing page is calculator-driven and returned no static figures across three attempts, including following the v2 documentation redirect. the widely-quoted 60 free minutes a month couldn't be confirmed on google's own page either, and google's docs point you to an internal model-comparison tool rather than publishing a headline accuracy figure.
one confirmed improvement worth noting: billing granularity moved from 15-second to 1-second rounding, which materially helps anyone transcribing lots of short clips. diarization is an available request parameter, though whether it carries a surcharge wasn't confirmable.
pros
+chirp 2 and chirp 3 models with broad language coverage
+billing now rounds to the second, not 15 seconds
+dedicated on-premise product line
+deepest integration for existing gcp users
cons
−no price figure extractable from its own page
−free-tier allowance unconfirmed on google's pages
every price is converted to usd per hour of audio, with the arithmetic shown in the entry, because no two vendors quote the same unit. where batch and real-time differ we give both, and where a headline rate requires a volume commitment we quote the pay-as-you-go rate instead and note the committed price separately.
accuracy claims are reported as what they are. a vendor's own word error rate is labelled as the vendor's, with the dataset named where one is given. we deliberately exclude vendor-published comparisons against rivals — several vendors publish tables of their competitors' error rates, and those are marketing rather than measurement.
where an independent leaderboard has evaluated a model we cite it separately. elevenlabs' scribe v2 ranks first for word error rate on artificial analysis's aa-agenttalk set at 1.5%, and the hugging face open asr leaderboard — which only scores open-weight models — is currently led by cohere transcribe and ibm granite speech rather than by whisper, which is no longer state of the art despite remaining the default baseline.
the training-data policy is treated as a pricing field rather than a privacy footnote, because at deepgram it literally is one. we report the default, whether opting out is possible, and what opting out costs.
whisper is the reference point for what managed transcription is worth: the same model that openai bills at $0.36 an hour is mit-licensed and free to run on your own hardware. every managed price on this page should be read against that.