verifier.org

Replicate

bills most models by how long they run, publishes token prices for two, and quotes them in different units to everyone else.

$3.75 / 1m in#one-off-runs#community-checkpoints-th

a good platform for the long tail of models, and close to the wrong tool for serving a current open-weight llm at volume.

replicate's own pricing page says most models are billed by the time they take to run, not by token. that suits image models, audio models and obscure community checkpoints — which is what it is genuinely good at — and makes it structurally incomparable to everything above it here.

only two entries carry per-token rates, and one of them, claude 3.7 sonnet, isn't an open-weight model at all. the other is deepseek r1 at $3.75 in and $10.00 out — a generation behind the v4 that most of this list serves, at several times the price. the output rate is printed as $0.01 per thousand tokens, a different unit convention from every other vendor here, which is an easy way to misread it by a factor of a thousand.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
$3.75 / 1m in
our verdict

a good platform for the long tail of models, and close to the wrong tool for serving a current open-weight llm at volume.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more llm inference providers

Anyscale

ray-based platform for serving and scaling your own models

#ray#self-serve

Hyperbolic

open-model inference and rented gpus on one account

#open-models#gpu

Parasail

serverless and dedicated endpoints for open-weight models

#serverless#dedicated

Novita AI

we checked this$0.135 / 1m in

near-cheapest prices with the full frontier catalogue behind them, and a batch discount the cheap rivals don't offer.

#most-teams-it-s-within-a#it-actually-carries-what