Replicate
bills most models by how long they run, publishes token prices for two, and quotes them in different units to everyone else.
a good platform for the long tail of models, and close to the wrong tool for serving a current open-weight llm at volume.
replicate's own pricing page says most models are billed by the time they take to run, not by token. that suits image models, audio models and obscure community checkpoints — which is what it is genuinely good at — and makes it structurally incomparable to everything above it here.
only two entries carry per-token rates, and one of them, claude 3.7 sonnet, isn't an open-weight model at all. the other is deepseek r1 at $3.75 in and $10.00 out — a generation behind the v4 that most of this list serves, at several times the price. the output rate is printed as $0.01 per thousand tokens, a different unit convention from every other vendor here, which is an easy way to misread it by a factor of a thousand.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- llm inference providers
- pricing
- $3.75 / 1m in
- website
- replicate.com
a good platform for the long tail of models, and close to the wrong tool for serving a current open-weight llm at volume.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.