verifier.org

best 14 llm inference providers

who to buy open-weight model tokens from, ranked on the price of the model you actually want, what they'll admit that price is, and whether they carry the model at all.

last reviewed 26 jul 2026 · 14 tools tested ·list curated by Onur Ozcanxin

the short version
best overallNovita AImost teams — it's within a rounding error of the cheapest and it actually carries what you'll want next quarter.90/100runner-upDeepInfrahigh-volume workloads on llama and deepseek where the model is already decided and the bill is the problem.89/100best free optionFireworks AIteams who want the frontier open models with the tiering and caching spelled out before they commit.88/100

the finding that reorganised this ranking: on the newest frontier open models, shopping around does almost nothing. glm-5.2 costs exactly $1.40 per million in and $4.40 out at fireworks, at baseten, at novita and at together — four independent vendors, the same two numbers. deepseek v4 pro is $1.74 and $3.48 at three of them. there is a reference price for a hot open model and most of the market simply prints it.

the spread is all in the older and smaller models, and it is enormous. llama 3.3 70b costs $0.10 per million input tokens at deepinfra, $0.135 at novita, $0.59 at groq, $0.60 at sambanova and $1.04 at together — a tenfold range for identical weights. if your workload runs on anything other than this quarter's flagship, provider choice is most of your bill.

the second axis is whether the vendor will tell you the price at all. cerebras builds the fastest inference hardware on earth and its per-model rates would not render for us across three attempts. nebius is mid-rebrand to 'token factory' with its price table behind javascript. hyperbolic's pricing page now redirects into a 404 and is not ranked below. featherless publishes three different model counts on three of its own pages.

the third is that 'per token' is no longer one thing. cached input runs 10-20% of standard at several vendors, deepinfra sells the same model at 0.8x, 1x and 1.5x depending on how long you'll wait, and fireworks splits every model into standard and priority. a headline rate now tells you less about your invoice than your cache hit rate does.

advertisement
  1. 1

    Novita AI

    near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

    90/100

    verdictthe best combination on this list: deepinfra's pricing without deepinfra's catalogue gaps, plus an explicit half-price batch tier.

    best for
    most teams — it's within a rounding error of the cheapest and it actually carries what you'll want next quarter.
    price
    $0.135 / 1m in
    pricing note
    llama 3.3 70b, $0.40 out; deepseek v4 pro $1.60/$3.20; glm-5.2 $1.40/$4.40; qwen3.7-max $1.25/$3.75
    free tier
    yes
    billing
    per-token
    llama 70b input
    $0.135/1m
    deepseek v4
    $1.60/1m in
    glm-5.2
    $1.40/1m in
    batch discount
    50%

    on the anchor model novita is $0.135 per million in — 35% above the cheapest price in the category and a seventh of the most expensive — and it undercuts the reference price on deepseek v4 pro at $1.60/$3.20 where three larger vendors all charge $1.74/$3.48. on glm-5.2 it prints the same $1.40/$4.40 as everyone else.

    the catalogue is the reason it ranks first rather than second. multiple deepseek generations, multiple glm generations including 5.2, the qwen family and llama — so the model you standardise on today and the one you migrate to in october are both here. deepinfra is cheaper and carries no glm at all, which is a worse trade than $0.035 per million.

    two extras that matter operationally: a stated 50% batch inference discount, and cache reads at 10-20% of the standard input rate. two models are served free outright. the gap is the usual one — no published rate limits and no speed figures, so capacity planning is guesswork until you're running.

    pros
    • +undercuts the reference price on deepseek v4 pro
    • +carries glm, deepseek, qwen and llama families in current versions
    • +explicit 50% batch discount and deep cache-read discounts
    • +two models served free
    cons
    • no published rate limits
    • no speed or throughput figures of its own
    • qwen3-max is 'tiered pricing' with no public rate
  2. 2

    DeepInfra

    the cheapest tokens in the category by a distance — with no glm models at all.

    89/100

    verdictunbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.

    best for
    high-volume workloads on llama and deepseek where the model is already decided and the bill is the problem.
    price
    $0.10 / 1m in
    pricing note
    llama 3.3 70b turbo, $0.32 out; llama 4 maverick $0.20/$0.80; deepseek v4 pro $1.30/$2.60
    free tier
    no
    billing
    per-token, 3 tiers
    llama 70b input
    $0.10/1m
    deepseek v4
    $1.30/1m in
    glm-5.2
    not served
    batch discount
    flex tier at 0.8x

    $0.10 per million input tokens on llama 3.3 70b is the lowest verified figure in this category, and it is not close: together charges $1.04 for the identical model. deepseek v4 pro at $1.30/$2.60 also undercuts the $1.74/$3.48 that fireworks, baseten and together all charge, so the discounting extends to the frontier rather than stopping at legacy models.

    the service-tier system is the cleverest pricing here and the most easily missed. the same model bills at 1x on standard, 1.5x on priority, and 0.8x on flex if you can tolerate latency — a batch discount in everything but name, applied per request rather than per job.

    the catalogue gap is real and specific: we found no glm model anywhere on deepinfra's pricing page. given glm-5.2's standing among current open models that is a significant hole, and it is the only reason this isn't first. dedicated gpu deployments run $0.89-$4.89/hr for teams that outgrow serverless.

    pros
    • +cheapest verified per-token prices in the category, by 10x on llama
    • +undercuts the reference price on deepseek v4 pro too
    • +flex tier at 0.8x for latency-tolerant work
    • +wide deepseek lineup across v3, v3.1, v3.2 and v4
    cons
    • no glm models found on the pricing page at all
    • qwen3.7-max is expensive here at $2.50/$7.50
    • no published rate limits or speed figures
  3. 3

    Fireworks AI

    the most legible pricing in the category, and the cheapest route to a frontier deepseek by a wide margin.

    88/100

    verdictpublishes more of its own pricing structure than anyone here, and deepseek v4 flash at $0.14/$0.28 is the standout value in the whole category.

    best for
    teams who want the frontier open models with the tiering and caching spelled out before they commit.
    price
    $0.14 / 1m in
    pricing note
    deepseek v4 flash, $0.28 out; v4 pro $1.74/$3.48; glm-5.2 $1.40/$4.40; qwen 3.7 plus $0.40/$1.60
    free tier
    yes
    billing
    per-token, std/priority
    llama 70b input
    generic tier, ~$0.90
    deepseek v4
    $1.74 pro / $0.14 flash
    glm-5.2
    $1.40/1m in
    batch discount
    50%

    the pricing docs are unusually honest about complexity instead of hiding it: every model has an explicit standard and priority rate, cached input is priced separately and steeply discounted — deepseek v4 pro's cached input is $0.145 against $1.74 standard — and models without a named row fall into published parameter-count tiers rather than a quote form.

    deepseek v4 flash is the find. at $0.14 in and $0.28 out it costs a twelfth of v4 pro, and for the large share of workloads that don't need the pro tier that is the cheapest frontier-adjacent capability on this page.

    it charges the reference $1.40/$4.40 on glm-5.2 like everyone else, adds 50% off batch inference, and rents on-demand gpus at $7/hr for h100 and h200 through $12/hr for b300 if you outgrow serverless. the one gap: no named llama flagship row, which likely means llama lands in the generic over-16b tier at $0.90 — pricier than the specialists.

    pros
    • +deepseek v4 flash at $0.14/$0.28 — best frontier value here
    • +standard vs priority and cached rates published per model
    • +unnamed models priced by parameter tier, not by quote
    • +50% batch discount plus on-demand gpu rental
    cons
    • no named llama flagship price — likely the generic $0.90 tier
    • openai-sdk compatibility not stated on its pricing pages
    • no published rate limits or speed figures
    advertisement
  4. 4

    Together AI

    every deployment shape from serverless to your own gpu cluster — at the highest llama price on this list.

    86/100

    verdictthe widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.

    best for
    teams who expect to graduate from serverless to reserved throughput or dedicated hardware without changing vendor.
    price
    $1.04 / 1m in
    pricing note
    llama 3.3 70b, $1.04 out; deepseek v4 pro $1.74/$3.48; glm-5.2 $1.40/$4.40; qwen3.6-plus $0.50/$3.00
    free tier
    no
    billing
    per-token to dedicated
    llama 70b input
    $1.04/1m
    deepseek v4
    $1.74/1m in
    glm-5.2
    $1.40/1m in
    batch discount
    none; commit instead

    the differentiator is deployment range rather than price. together sells serverless tokens, provisioned throughput units with guaranteed capacity, single-tenant dedicated inference instances at $5.49-$8.99/hr, and reserved gpu clusters with volume discounts up to 32%. nobody else here covers that whole span, and the frontier open models land on it fast.

    on price it is the most expensive place to run the anchor model: $1.04 per million input for llama 3.3 70b, against $0.10 at deepinfra. it charges the reference rate on glm-5.2 and deepseek v4 pro like the rest of the field, and qwen3.6-plus at $0.50/$3.00 is genuinely competitive — so the premium is concentrated on llama rather than universal.

    watch two things. cached input is dramatically cheaper across deepseek, glm and qwen — $0.20 against $1.74 on v4 pro — so your effective rate depends on hit rate more than on the sticker. and there is no batch api discount; savings come from committing to capacity instead, which is a different financial decision.

    pros
    • +serverless through ptu, dedicated instances and full clusters
    • +same-day access to new frontier open releases
    • +steep cached-input rates across deepseek, glm and qwen
    • +volume discounts up to 32% on commitments
    cons
    • 10x deepinfra's price on the anchor llama model
    • no batch api discount — savings require capacity commitments
    • no published rate limits; dynamic and account-level
    • $4 minimum per fine-tuning job
  5. 5

    Groq

    the only vendor here that prints tokens-per-second next to the price — on a catalogue of five models.

    85/100

    verdictuniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.

    best for
    latency-sensitive products that can live on llama or gpt-oss and want the speed number in writing.
    price
    $0.59 / 1m in
    pricing note
    llama 3.3 70b versatile, $0.79 out, 394 tps claimed; qwen 3.6 27b $0.60/$3.00 at 500 tps
    free tier
    yes
    billing
    per-token
    llama 70b input
    $0.59/1m
    deepseek v4
    not served
    glm-5.2
    not served
    batch discount
    50%

    groq publishes a tokens-per-second figure per model directly on its pricing page: 394 tps for llama 3.3 70b, 500 for qwen 3.6 27b and gpt-oss 120b, 840 for llama 3.1 8b, 1,000 for gpt-oss 20b. every other vendor in this category either markets a vague multiple or says nothing. treating speed as a published spec rather than a claim is the right instinct and nobody else has it.

    the price is mid-field and reasonable — $0.59/$0.79 on the anchor model, roughly six times deepinfra but well under together — with a 50% batch discount, 50% off cached input, and drop-in openai sdk compatibility.

    the catalogue is the constraint. five models on the public pricing page, no deepseek and no glm anywhere, and minimax reserved for enterprise. the free tier is real but tight: 30 requests a minute and 1,000 a day on the 70b model, which is a demo allowance rather than a runway.

    pros
    • +per-model tokens-per-second published on the pricing page
    • +50% off both batch and cached input
    • +openai sdk drop-in compatibility
    • +genuinely low latency at a mid-field price
    cons
    • only five models on the public pricing page
    • no deepseek and no glm models at all
    • free tier capped at 1,000 requests a day on the 70b
    • minimax restricted to enterprise customers
  6. 6

    OpenRouter

    one endpoint in front of the whole market, at genuinely no markup — which makes the routing itself the product.

    84/100

    verdictcharges nothing to sit in the middle, which is rare enough to be worth using — you just can't look up 'the openrouter price' for anything, because there isn't one.

    best for
    anyone still choosing, still comparing, or wanting to switch providers without shipping a code change.
    price
    the routed provider's price
    pricing note
    pass-through pricing with no stated markup on standard inference; byok is free to 1m requests/month then 5% of the equivalent openrouter cost
    free tier
    yes
    billing
    pass-through
    llama 70b input
    routed provider's rate
    deepseek v4
    via routed provider
    glm-5.2
    via routed provider
    batch discount
    provider's own

    openrouter's own faq states it passes through the underlying provider's pricing without markup on standard inference. for an aggregator that is a genuinely good deal, and it removes the usual objection — our ai image models ranking documents aggregators charging 12-25% over going direct.

    what you buy is optionality. one openai-compatible endpoint reaches most of the vendors ranked above, so switching provider or model is configuration rather than engineering, and a free tier lets you try models at 50 requests a day, rising to 1,000 once you've bought $10 of credit.

    two caveats. the price of any given call is whatever the routed provider charges, so there is no openrouter price list to plan against — you are still doing the comparison this page exists to help with. and bring-your-own-key is free only to a million requests a month, after which it costs 5% of what openrouter would have charged.

    pros
    • +no markup on standard inference, per its own faq
    • +one openai-compatible endpoint across most of this list
    • +switching provider or model needs no code change
    • +free tier for evaluation, 1,000 requests/day after $10 credit
    cons
    • no price list of its own — cost depends entirely on routing
    • byok costs 5% beyond 1m requests a month
    • free models explicitly not recommended for production
    • model catalogue renders only with javascript
  7. 7

    Baseten

    the compliance-first option — soc 2 type ii and hipaa, hybrid deployment, and no public price for half its catalogue.

    81/100

    verdictthe right answer when compliance drives the decision — priced at the market reference on what it does publish, and silent on the rest.

    best for
    regulated teams who need the certifications and the option to run the same stack in their own environment.
    price
    $0.60 / 1m in
    pricing note
    glm 4.7, $2.20 out; glm-5.2 $1.40/$4.40; deepseek v4 $1.74/$3.48; llama and qwen unpriced publicly
    free tier
    yes
    billing
    per-token + dedicated
    llama 70b input
    not published
    deepseek v4
    $1.74/1m in
    glm-5.2
    $1.40/1m in
    batch discount
    none found

    baseten is the only entry here leading with soc 2 type ii and hipaa, and one of the few offering self-hosted and hybrid deployment alongside its cloud. for teams whose blocker is a security review rather than a price, that combination decides it.

    published pricing sits exactly on the market reference — glm-5.2 at $1.40/$4.40, deepseek v4 at $1.74/$3.48 — with glm 4.7 at $0.60/$2.20 as a cheaper step down and a 'fast' glm-5.2 variant at $2.10/$6.60 for latency-sensitive work. billing is compute-time only, with no idle charge.

    the frustration is that only nine models carry a public per-token rate. llama 4, llama 3.3, qwen3.6, kimi k2.7 and nemotron all appear in the model library with no price beside them, which in practice means dedicated deployment and a conversation. that's a defensible enterprise motion and a bad experience if you just wanted to know what llama costs.

    pros
    • +soc 2 type ii and hipaa, rare in this field
    • +cloud, self-hosted and hybrid deployment
    • +compute-time billing with no idle charges
    • +broad library including kimi k2.7 and nemotron 3 ultra
    cons
    • llama and qwen models carry no public per-token price
    • only nine models priced on the public page
    • no batch discount found
    • no published rate limits or speed figures
  8. 8

    Cerebras

    the fastest inference hardware built, attached to a pricing page that wouldn't tell us what anything costs.

    76/100

    verdictthe technology is genuinely in a class of its own and the subscription tiers are good value — but we could not verify a single per-token rate, and that costs it eight places.

    best for
    developers who want extreme throughput and can work within a daily-token subscription rather than a per-token budget.
    price
    unpublished
    pricing note
    per-model token rates did not render across three fetch attempts; subscription plans are $50/mo for up to 24m tokens/day and $200/mo for up to 120m
    free tier
    yes
    billing
    per-token + subscriptions
    llama 70b input
    unpublished
    deepseek v4
    unpublished
    glm-5.2
    unpublished
    batch discount
    unpublished

    the wafer-scale hardware is not marketing: independent leaderboards have put cerebras far ahead of gpu-based inference on large models, and no competitor here claims comparable throughput. if generation speed is the product constraint, this is the shortlist.

    what we could verify is thin. $5 of free credits, a $10 self-serve minimum, and two subscription plans — code pro at $50/month for up to 24 million tokens a day, code max at $200/month for up to 120 million. those are strong allowances if your usage fits the daily cap shape.

    what we could not verify is the per-model, per-token price, across three attempts at the vendor's own pricing page. the site's only speed claim in readable text is '20x faster than openai and anthropic', with no model attached. a vendor selling on performance that publishes neither a numeric benchmark nor a token rate in machine-readable form is asking buyers to take both on faith. one preview model also carries a deprecation date of 17 august 2026.

    pros
    • +fastest inference hardware in the category by a wide margin
    • +generous daily token allowances on flat subscriptions
    • +$5 free credits and a low $10 self-serve minimum
    cons
    • per-model token prices unreadable across three attempts
    • speed marketed as '20x' with no model or figure attached
    • no published numeric rate limits
    • a preview model is already scheduled for deprecation
  9. 9

    SambaNova

    custom silicon built for speed, from a vendor that publishes no speed figures and stopped at deepseek v3.2.

    73/100

    verdictfair llama pricing on interesting hardware, undone by a catalogue that has fallen a generation behind and documentation with holes in it.

    best for
    teams already committed to deepseek v3.x who want it on dataflow hardware rather than gpus.
    price
    $0.60 / 1m in
    pricing note
    meta-llama-3.3-70b-instruct, $1.20 out; deepseek v3.1 and v3.2 both $3.00/$4.50
    free tier
    no
    billing
    per-token
    llama 70b input
    $0.60/1m
    deepseek v4
    not served; v3.2 only
    glm-5.2
    not served
    batch discount
    not published

    the llama price is competitive — $0.60/$1.20 on the anchor model, roughly matching groq — and the rdu architecture is a genuine third approach alongside gpus and groq's lpu. openai-sdk compatibility is confirmed in its docs.

    the catalogue is where it falls behind. six models, no glm, no qwen, and deepseek stops at v3.1 and v3.2 while most of this list has moved to v4 — both v3 variants priced at $3.00/$4.50, which is more than fireworks charges for v4 pro. a cached-input rate on minimax at $0.06 against $0.60 standard is the one striking discount.

    the documentation gaps compound it. no tokens-per-second figures anywhere, despite speed being the company's entire hardware thesis, and the rate-limits page returned a 404 when we looked. for a vendor asking developers to adopt unfamiliar silicon, that is the wrong pair of things to leave blank.

    pros
    • +competitive llama 3.3 70b pricing at $0.60/$1.20
    • +genuinely distinct rdu hardware architecture
    • +openai-sdk compatibility confirmed in its docs
    • +steep cached-input discount on minimax
    cons
    • deepseek stops at v3.2 — no v4 at all
    • no glm or qwen models on the pricing page
    • publishes no speed figures despite selling speed
    • rate-limits documentation returned 404
  10. 10

    Nebius AI Studio

    sixty-plus current open models behind a price table that won't render, during a rename.

    70/100

    verdictthe catalogue looks right and the european base may matter for your data rules — but you cannot compare it on price without signing up, and it's mid-rebrand.

    best for
    buyers who will open the dashboard themselves and want breadth from a european provider.
    price
    unpublished
    pricing note
    per-model rates live in a javascript app at tokenfactory.nebius.com that returned no readable pricing; deepseek v4 pro, glm-5.1 and qwen3 are confirmed as offered
    free tier
    no
    billing
    per-token
    llama 70b input
    unpublished
    deepseek v4
    offered, price unreadable
    glm-5.2
    glm-5.1 listed
    batch discount
    unpublished

    sixty-plus open-source models is real breadth, and the named ones — deepseek v4 pro, glm-5.1, qwen3 variants, llama — are the current generation rather than a stale list. the api is openai-compatible, and a european provider is a genuine consideration for teams with data-residency constraints.

    the pricing is inaccessible. the per-model table lives at a javascript-rendered dashboard url, and no rates could be read from it. we will not publish figures for a vendor whose own price list we couldn't open, so every rate here is unverified.

    it is also mid-rebrand, from 'nebius ai studio' to 'nebius token factory', with both names live on its own properties. that is a small thing, but combined with unreadable pricing and no published rate limits it makes evaluation harder than it should be for a provider whose catalogue deserves a look.

    pros
    • +60+ open models, current generation
    • +openai-compatible api
    • +european provider — relevant for data residency
    cons
    • per-model prices unreadable without javascript
    • no published rate limits or free-tier detail
    • mid-rebrand with two product names live at once
    • glm listed at 5.1 rather than 5.2
  11. 11

    inference.net

    the cheapest llama 4 scout price we found, wrapped in a plan structure that meters something other than tokens.

    68/100

    verdictthe headline prices are excellent and the catalogue we could verify is three models deep — a good deal you can't fully evaluate.

    best for
    cost-driven workloads on llama 4 scout or maverick that fit inside the request allowances.
    price
    $0.08 / 1m in
    pricing note
    llama 4 scout, $0.15 out; llama 4 maverick $0.35/$0.40; deepseek v3.2 $0.14/$0.28
    free tier
    yes
    billing
    per-token + request plans
    llama 70b input
    llama 4 only, $0.08
    deepseek v4
    not found; v3.2 $0.14
    glm-5.2
    named, not priced
    batch discount
    not published

    llama 4 scout at $0.08/$0.15 and maverick at $0.35/$0.40 are the cheapest rates we verified for those models anywhere in this category, and deepseek v3.2 at $0.14/$0.28 is likewise strong.

    the plan layer is unusual and needs reading twice. on top of per-token usage sit tiers metered by 'gateway requests': pay-as-you-go includes a million a month at 30 requests per minute, growth costs $250/month for fifty million at 250 per minute. that request ceiling, not the token price, is what will actually constrain a busy application.

    verification thinned out fast. the main pricing table didn't render as static text, so the figures above came from the vendor's own comparison post rather than its price list, and we found no published rate for glm-5.2, deepseek v4 or any qwen model despite glm-5.2 being named elsewhere on the site. total catalogue size is unpublished.

    pros
    • +cheapest verified llama 4 scout and maverick prices here
    • +a million gateway requests included at the free entry tier
    • +clear per-minute request limits, which most rivals don't publish
    cons
    • main price table doesn't render as static text
    • no verifiable price for glm-5.2, deepseek v4 or qwen
    • catalogue size unpublished
    • 30 requests/minute on the entry tier is restrictive
  12. 12

    Featherless AI

    a flat monthly fee for unlimited tokens across tens of thousands of models — and three different counts of how many, on its own pages.

    65/100

    verdictgenuinely different economics that will beat per-token pricing for some workloads — sold by a vendor that can't keep its own catalogue claims straight.

    best for
    heavy, steady usage on small and mid-sized models where per-token billing would run away from you.
    price
    $25/mo
    pricing note
    premium: unlimited tokens, 4 concurrent units, catalogue models; agent max at $200/mo is the tier that unlocks deepseek, kimi and glm
    free tier
    no
    billing
    flat subscription
    llama 70b input
    unmetered on plan
    deepseek v4
    $200/mo tier
    glm-5.2
    $200/mo tier
    batch discount
    n/a

    the model is the point: a flat monthly fee for unlimited tokens, gated by concurrency rather than volume. $25 buys four concurrent units against the catalogue; $100 covers models up to 229b with eight units and 256k context; $200 unlocks anything including deepseek, kimi and glm. for constant high-volume traffic on smaller models, that inverts the usual arithmetic in your favour.

    the catalogue breadth is the other pitch, and it is where the credibility wobbles. featherless states '47,300+ models' on one page, '20,000+' on another, and '6,700+' in its own blog. we cannot tell you which is true, and a vendor whose headline number varies sevenfold across its own properties is asking to be taken on trust.

    two structural limits: the flagship open models everyone else prices at $1.40 per million sit behind the $200/month tier here, and there is no per-token option for like-for-like comparison unless you use the separate per-request product. openai compatibility isn't confirmed on the pages we could read.

    pros
    • +flat fee for unlimited tokens — beats per-token at steady high volume
    • +very large catalogue of small and mid-sized open models
    • +concurrency-based capacity is predictable to plan around
    cons
    • its own pages claim 6,700+, 20,000+ and 47,300+ models
    • frontier models require the $200/month tier
    • no free tier stated
    • openai-sdk compatibility unconfirmed
  13. 14

    Replicate

    bills most models by how long they run, publishes token prices for two, and quotes them in different units to everyone else.

    58/100

    verdicta good platform for the long tail of models, and close to the wrong tool for serving a current open-weight llm at volume.

    best for
    one-off runs and community checkpoints that no managed api carries.
    price
    $3.75 / 1m in
    pricing note
    deepseek r1, $10.00 out — printed on the page as $0.01 per thousand output tokens; most other models bill per second of compute
    free tier
    no
    billing
    per-run compute
    llama 70b input
    not priced
    deepseek v4
    not served; r1 only
    glm-5.2
    not served
    batch discount
    not published

    replicate's own pricing page says most models are billed by the time they take to run, not by token. that suits image models, audio models and obscure community checkpoints — which is what it is genuinely good at — and makes it structurally incomparable to everything above it here.

    only two entries carry per-token rates, and one of them, claude 3.7 sonnet, isn't an open-weight model at all. the other is deepseek r1 at $3.75 in and $10.00 out — a generation behind the v4 that most of this list serves, at several times the price. the output rate is printed as $0.01 per thousand tokens, a different unit convention from every other vendor here, which is an easy way to misread it by a factor of a thousand.

    we tried to price a current llama flagship on it and the model page returned a 404 within our budget. if you want llama, glm, qwen or deepseek v4 behind an api, everything above this entry does it more cheaply and more legibly.

    pros
    • +unmatched breadth of community and custom checkpoints
    • +per-second billing suits irregular, one-off runs
    • +strong for image, audio and video models
    cons
    • most models billed by run time, not tokens
    • only deepseek r1 among open llms has a published token rate
    • output priced per thousand tokens while rivals quote per million
    • current llama flagship page returned 404

how this ranking was made

prices are the vendor's own published per-million-token rates for named models, read from the vendor's pricing page or pricing docs on the review date. we quote input and output separately because the ratio between them varies by more than the absolute numbers do.

llama 3.3 70b is the comparison anchor wherever a vendor serves it. it is not the best model here — it is the one nearly everybody carries, which makes it the only honest like-for-like column. where a vendor doesn't serve it we say so rather than substituting something cheaper.

we rank on published, standard-tier, uncached prices. cached-input and batch discounts are described in the entry because they are real, but a rate you only get on a cache hit is not a price you can plan with.

model availability is treated as a first-class fact. deepinfra has no glm at all; sambanova stops at deepseek v3.2; baseten lists llama and qwen in its library but publishes no per-token rate for them. a cheap provider that doesn't carry your model is not cheap.

hyperbolic is not ranked. its pricing url redirects to a page that returns 404, so there is no figure on it we could verify. we would rather leave a gap than publish a number from a comparison blog.

three entries are not per-token products at all — modal sells gpu seconds, replicate bills most models by run time, featherless sells a flat monthly subscription. they are ranked at the bottom with an explanation rather than omitted, because buyers keep finding them in these comparisons and deserve to know why the maths doesn't line up.

we publish no independent speed measurements. groq is the only vendor here that prints per-model tokens-per-second on its own pricing page, and we report those as its claims rather than as verified figures.

our general methodology and disclosures →
was this useful?