verifier.org

Together AI

every deployment shape from serverless to your own gpu cluster — at the highest llama price on this list.

$1.04 / 1m in#teams-who-expect-to-grad

the widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.

the differentiator is deployment range rather than price. together sells serverless tokens, provisioned throughput units with guaranteed capacity, single-tenant dedicated inference instances at $5.49-$8.99/hr, and reserved gpu clusters with volume discounts up to 32%. nobody else here covers that whole span, and the frontier open models land on it fast.

on price it is the most expensive place to run the anchor model: $1.04 per million input for llama 3.3 70b, against $0.10 at deepinfra. it charges the reference rate on glm-5.2 and deepseek v4 pro like the rest of the field, and qwen3.6-plus at $0.50/$3.00 is genuinely competitive — so the premium is concentrated on llama rather than universal.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$1.04 / 1m in
our verdict

the widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more llm inference providers

Anyscale

ray-based platform for serving and scaling your own models

#ray#self-serve

Hyperbolic

open-model inference and rented gpus on one account

#open-models#gpu

Parasail

serverless and dedicated endpoints for open-weight models

#serverless#dedicated

Novita AI

we checked this$0.135 / 1m in

near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

#most-teams-it-s-within-a#it-actually-carries-what