DeepInfra
the cheapest tokens in the category by a distance — with no glm models at all.
unbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.
$0.10 per million input tokens on llama 3.3 70b is the lowest verified figure in this category, and it is not close: together charges $1.04 for the identical model. deepseek v4 pro at $1.30/$2.60 also undercuts the $1.74/$3.48 that fireworks, baseten and together all charge, so the discounting extends to the frontier rather than stopping at legacy models.
the service-tier system is the cleverest pricing here and the most easily missed. the same model bills at 1x on standard, 1.5x on priority, and 0.8x on flex if you can tolerate latency — a batch discount in everything but name, applied per request rather than per job.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- llm inference providers
- pricing
- $0.10 / 1m in
- website
- deepinfra.com
unbeatable on price — a tenth of together's rate for the same llama weights — provided nothing you need is a glm model.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.