Modal
gpu seconds, not model tokens — the right tool for a different job than this page is about.
excellent serverless gpu with no idle billing, ranked low only because it isn't the product this category compares: there is no token price, because you build the endpoint.
modal sells raw serverless gpu by the second: h100 at $0.001097, a100 80gb at $0.000694, b300 at $0.001972, down to t4 at $0.000164, billed only while running with no idle charge. the starter plan includes $30 of monthly credits and ten concurrent gpus.
there is no hosted open-weight model api here and no per-token rate to compare, because the serving layer is yours — you deploy vllm or sglang and operate it. that is genuine freedom for a custom fine-tune or an unusual checkpoint nobody else hosts, and pure overhead if you just wanted glm-5.2 behind an endpoint.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- llm inference providers
- pricing
- $0.001097/s
- website
- modal.com
excellent serverless gpu with no idle billing, ranked low only because it isn't the product this category compares: there is no token price, because you build the endpoint.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.