DeepInfra ranks #2 of 14 in our llm inference providers testing. the cheapest tokens in the category by a distance — with no glm models at all..
89/100
unbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.
why people look for an alternative
−no glm models found on the pricing page at all
−qwen3.7-max is expensive here at $2.50/$7.50
−no published rate limits or speed figures
stay with DeepInfra if cheapest verified per-token prices in the category, by 10x on llama is the thing you care about most — nothing below beats it on that.
#4 in llm inference providers · every deployment shape from serverless to your own gpu cluster — at the highest llama price on this list.
86/100
verdictthe widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.
Together AI vs DeepInfra
DeepInfra
Together AI
price
$0.10 / 1m in
$1.04 / 1m in
free tier
no
no
billing
per-token, 3 tiers
per-token to dedicated
llama 70b input
$0.10/1m
$1.04/1m
deepseek v4
$1.30/1m in
$1.74/1m in
glm-5.2
not served
$1.40/1m in
batch discount
flex tier at 0.8x
none; commit instead
switch forteams who expect to graduate from serverless to reserved throughput or dedicated hardware without changing vendor.
pros
+serverless through ptu, dedicated instances and full clusters
+same-day access to new frontier open releases
+steep cached-input rates across deepseek, glm and qwen
+volume discounts up to 32% on commitments
cons
−10x deepinfra's price on the anchor llama model
−no batch api discount — savings require capacity commitments
−no published rate limits; dynamic and account-level
#5 in llm inference providers · the only vendor here that prints tokens-per-second next to the price — on a catalogue of five models.
85/100
verdictuniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.
Groq vs DeepInfra
DeepInfra
Groq
price
$0.10 / 1m in
$0.59 / 1m in
free tier
no
yes
billing
per-token, 3 tiers
per-token
llama 70b input
$0.10/1m
$0.59/1m
deepseek v4
$1.30/1m in
not served
glm-5.2
not served
not served
batch discount
flex tier at 0.8x
50%
switch forlatency-sensitive products that can live on llama or gpt-oss and want the speed number in writing.
pros
+per-model tokens-per-second published on the pricing page
+50% off both batch and cached input
+openai sdk drop-in compatibility
+genuinely low latency at a mid-field price
cons
−only five models on the public pricing page
−no deepseek and no glm models at all
−free tier capped at 1,000 requests a day on the 70b
#6 in llm inference providers · one endpoint in front of the whole market, at genuinely no markup — which makes the routing itself the product.
84/100
verdictcharges nothing to sit in the middle, which is rare enough to be worth using — you just can't look up 'the openrouter price' for anything, because there isn't one.
OpenRouter vs DeepInfra
DeepInfra
OpenRouter
price
$0.10 / 1m in
the routed provider's price
free tier
no
yes
billing
per-token, 3 tiers
pass-through
llama 70b input
$0.10/1m
routed provider's rate
deepseek v4
$1.30/1m in
via routed provider
glm-5.2
not served
via routed provider
batch discount
flex tier at 0.8x
provider's own
switch foranyone still choosing, still comparing, or wanting to switch providers without shipping a code change.
pros
+no markup on standard inference, per its own faq
+one openai-compatible endpoint across most of this list
+switching provider or model needs no code change
+free tier for evaluation, 1,000 requests/day after $10 credit
cons
−no price list of its own — cost depends entirely on routing
−byok costs 5% beyond 1m requests a month
−free models explicitly not recommended for production
#8 in llm inference providers · the fastest inference hardware built, attached to a pricing page that wouldn't tell us what anything costs.
76/100
verdictthe technology is genuinely in a class of its own and the subscription tiers are good value — but we could not verify a single per-token rate, and that costs it eight places.
Cerebras vs DeepInfra
DeepInfra
Cerebras
price
$0.10 / 1m in
unpublished
free tier
no
yes
billing
per-token, 3 tiers
per-token + subscriptions
llama 70b input
$0.10/1m
unpublished
deepseek v4
$1.30/1m in
unpublished
glm-5.2
not served
unpublished
batch discount
flex tier at 0.8x
unpublished
switch fordevelopers who want extreme throughput and can work within a daily-token subscription rather than a per-token budget.
pros
+fastest inference hardware in the category by a wide margin
+generous daily token allowances on flat subscriptions
+$5 free credits and a low $10 self-serve minimum
cons
−per-model token prices unreadable across three attempts
−speed marketed as '20x' with no model or figure attached
−no published numeric rate limits
−a preview model is already scheduled for deprecation
#10 in llm inference providers · sixty-plus current open models behind a price table that won't render, during a rename.
70/100
verdictthe catalogue looks right and the european base may matter for your data rules — but you cannot compare it on price without signing up, and it's mid-rebrand.
Nebius AI Studio vs DeepInfra
DeepInfra
Nebius AI Studio
price
$0.10 / 1m in
unpublished
free tier
no
no
billing
per-token, 3 tiers
per-token
llama 70b input
$0.10/1m
unpublished
deepseek v4
$1.30/1m in
offered, price unreadable
glm-5.2
not served
glm-5.1 listed
batch discount
flex tier at 0.8x
unpublished
switch forbuyers who will open the dashboard themselves and want breadth from a european provider.
every tool on this page went through the same test as DeepInfra — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.