verifier.org

SmolLM3-3B

the fully-open lineage pick — apache 2.0 from hugging face, with think and no-think modes.

free (apache 2.0)#instruction-following#tool-use-at-3b#anyone-who-values-a-tran

punches above its size on instruction following and comes from the one vendor here whose whole reason for existing is openness — with weaker long-context recall than the spec suggests.

76.7% on ifeval against qwen2.5-3b's 65.6% is a wide margin for instruction following at this size, and the dual think and no-think modes let you spend reasoning tokens only when a task warrants it. at roughly 2gb quantised it runs on anything, cpu included.

apache 2.0 with no conditions, from hugging face, which has a stronger institutional commitment to publishing training details than any commercial vendor on this page.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free (apache 2.0)
our verdict

punches above its size on instruction following and comes from the one vendor here whose whole reason for existing is openness — with weaker long-context recall than the spec suggests.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more small on-device llms

TinyLlama

1.1b model trained on 3t tokens, a common baseline for tiny deployments

#baseline#1b

MobileLLM

meta's research line for sub-billion-parameter on-device models

#research#sub-1b

H2O Danube

h2o.ai's small models released for edge and offline use

#edge

Gemma 4 (E2B / E4B / 12B)

we checked thisfree (apache 2.0)

text, images and audio on a phone, under an apache licence google took three generations to arrive at.

#anything-that-needs-to-s#on-hardware-you-already-