verifier.org

Groq

the only vendor here that prints tokens-per-second next to the price — on a catalogue of five models.

$0.59 / 1m in#latency-sensitive-produc#want-the-speed-number-in

uniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.

groq publishes a tokens-per-second figure per model directly on its pricing page: 394 tps for llama 3.3 70b, 500 for qwen 3.6 27b and gpt-oss 120b, 840 for llama 3.1 8b, 1,000 for gpt-oss 20b. every other vendor in this category either markets a vague multiple or says nothing. treating speed as a published spec rather than a claim is the right instinct and nobody else has it.

the price is mid-field and reasonable — $0.59/$0.79 on the anchor model, roughly six times deepinfra but well under together — with a 50% batch discount, 50% off cached input, and drop-in openai sdk compatibility.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$0.59 / 1m in
website
groq.com
our verdict

uniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more llm inference providers

Anyscale

ray-based platform for serving and scaling your own models

#ray#self-serve

Hyperbolic

open-model inference and rented gpus on one account

#open-models#gpu

Parasail

serverless and dedicated endpoints for open-weight models

#serverless#dedicated

Novita AI

we checked this$0.135 / 1m in

near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

#most-teams-it-s-within-a#it-actually-carries-what