verifier.org

DeepInfra

the cheapest tokens in the category by a distance — with no glm models at all.

$0.10 / 1m in#high-volume-workloads-on#deepseek-where-the-model#the-bill-is-the-problem

unbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.

$0.10 per million input tokens on llama 3.3 70b is the lowest verified figure in this category, and it is not close: together charges $1.04 for the identical model. deepseek v4 pro at $1.30/$2.60 also undercuts the $1.74/$3.48 that fireworks, baseten and together all charge, so the discounting extends to the frontier rather than stopping at legacy models.

the service-tier system is the cleverest pricing here and the most easily missed. the same model bills at 1x on standard, 1.5x on priority, and 0.8x on flex if you can tolerate latency — a batch discount in everything but name, applied per request rather than per job.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$0.10 / 1m in
our verdict

unbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more llm inference providers

Anyscale

ray-based platform for serving and scaling your own models

#ray#self-serve

Hyperbolic

open-model inference and rented gpus on one account

#open-models#gpu

Parasail

serverless and dedicated endpoints for open-weight models

#serverless#dedicated

Novita AI

we checked this$0.135 / 1m in

near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

#most-teams-it-s-within-a#it-actually-carries-what