verifier.org

Modal

gpu seconds, not model tokens — the right tool for a different job than this page is about.

$0.001097/s#teams-who-want-to-run-th#a-custom-checkpoint-with

excellent serverless gpu with no idle billing, ranked low only because it isn't the product this category compares: there is no token price, because you build the endpoint.

modal sells raw serverless gpu by the second: h100 at $0.001097, a100 80gb at $0.000694, b300 at $0.001972, down to t4 at $0.000164, billed only while running with no idle charge. the starter plan includes $30 of monthly credits and ten concurrent gpus.

there is no hosted open-weight model api here and no per-token rate to compare, because the serving layer is yours — you deploy vllm or sglang and operate it. that is genuine freedom for a custom fine-tune or an unusual checkpoint nobody else hosts, and pure overhead if you just wanted glm-5.2 behind an endpoint.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$0.001097/s
website
modal.com
our verdict

excellent serverless gpu with no idle billing, ranked low only because it isn't the product this category compares: there is no token price, because you build the endpoint.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more llm inference providers

Anyscale

ray-based platform for serving and scaling your own models

#ray#self-serve

Hyperbolic

open-model inference and rented gpus on one account

#open-models#gpu

Parasail

serverless and dedicated endpoints for open-weight models

#serverless#dedicated

Novita AI

we checked this$0.135 / 1m in

near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

#most-teams-it-s-within-a#it-actually-carries-what