Groq
the only vendor here that prints tokens-per-second next to the price — on a catalogue of five models.
uniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.
groq publishes a tokens-per-second figure per model directly on its pricing page: 394 tps for llama 3.3 70b, 500 for qwen 3.6 27b and gpt-oss 120b, 840 for llama 3.1 8b, 1,000 for gpt-oss 20b. every other vendor in this category either markets a vague multiple or says nothing. treating speed as a published spec rather than a claim is the right instinct and nobody else has it.
the price is mid-field and reasonable — $0.59/$0.79 on the anchor model, roughly six times deepinfra but well under together — with a 50% batch discount, 50% off cached input, and drop-in openai sdk compatibility.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- llm inference providers
- pricing
- $0.59 / 1m in
- website
- groq.com
uniquely transparent about throughput and priced fairly for it — but a five-model catalogue with no deepseek and no glm rules it out for a lot of work.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.