verifier.org

LFM2.5-8B-A1B

moe cleverness that runs like a 1.5b model — free only until your company makes $10m.

free under $10m revenue#small-companies#individuals-who-want-moe#will-stay-well-under-the

genuinely fast for its capacity and the only entry here with a revenue cap — phrased permissively enough that a growing company could sail past it without noticing.

the architecture is the appeal: 8.3b total parameters with only 1.5b active per token, so it generates at roughly the speed of a 1.5b model while carrying far more capacity. it runs across llama.cpp, mlx, vllm and sglang, so apple silicon gets a native path.

the memory maths is the thing people get wrong about moe, and it's worth stating plainly: all 8.3b parameters must be resident, about 5gb at four-bit. mixture-of-experts shrinks compute, not footprint. this is not a 1.5b model's memory profile.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free under $10m revenue
our verdict

genuinely fast for its capacity and the only entry here with a revenue cap — phrased permissively enough that a growing company could sail past it without noticing.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more small on-device llms

TinyLlama

1.1b model trained on 3t tokens, a common baseline for tiny deployments

#baseline#1b

MobileLLM

meta's research line for sub-billion-parameter on-device models

#research#sub-1b

H2O Danube

h2o.ai's small models released for edge and offline use

#edge

Gemma 4 (E2B / E4B / 12B)

we checked thisfree (apache 2.0)

text, images and audio on a phone, under an apache licence google took three generations to arrive at.

#anything-that-needs-to-s#on-hardware-you-already-