LFM2.5-8B-A1B
moe cleverness that runs like a 1.5b model — free only until your company makes $10m.
genuinely fast for its capacity and the only entry here with a revenue cap — phrased permissively enough that a growing company could sail past it without noticing.
the architecture is the appeal: 8.3b total parameters with only 1.5b active per token, so it generates at roughly the speed of a 1.5b model while carrying far more capacity. it runs across llama.cpp, mlx, vllm and sglang, so apple silicon gets a native path.
the memory maths is the thing people get wrong about moe, and it's worth stating plainly: all 8.3b parameters must be resident, about 5gb at four-bit. mixture-of-experts shrinks compute, not footprint. this is not a 1.5b model's memory profile.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- small on-device llms
- pricing
- free under $10m revenue
- website
- docs.liquid.ai
genuinely fast for its capacity and the only entry here with a revenue cap — phrased permissively enough that a growing company could sail past it without noticing.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.