Together AI
every deployment shape from serverless to your own gpu cluster — at the highest llama price on this list.
the widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.
the differentiator is deployment range rather than price. together sells serverless tokens, provisioned throughput units with guaranteed capacity, single-tenant dedicated inference instances at $5.49-$8.99/hr, and reserved gpu clusters with volume discounts up to 32%. nobody else here covers that whole span, and the frontier open models land on it fast.
on price it is the most expensive place to run the anchor model: $1.04 per million input for llama 3.3 70b, against $0.10 at deepinfra. it charges the reference rate on glm-5.2 and deepseek v4 pro like the rest of the field, and qwen3.6-plus at $0.50/$3.00 is genuinely competitive — so the premium is concentrated on llama rather than universal.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- llm inference providers
- pricing
- $1.04 / 1m in
- website
- together.ai
the widest ladder in the category — pay-per-token, provisioned throughput, dedicated instances, whole clusters — and you pay for the ladder on every token.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.