verifier.org

gpt-oss-20b

openai's small one — 21b of capacity in 16gb, with reasoning effort you can dial down.

free (apache 2.0)#agentic-tool-calling-on-#where-you-want-to-trade-

the most capable model here if you have the memory, with the genuinely useful ability to turn reasoning effort down when a task doesn't need it.

21b total parameters with 3.6b active, shipped post-trained in mxfp4 so the native format is already roughly four-bit, and openai states it fits within 16gb. that is the top of the consumer envelope rather than the middle, but it buys noticeably more capability than the 3-4b models around it.

configurable reasoning effort — low, medium, high — is the feature worth having. most small models make you choose a model per latency budget; this one lets you choose per request, which matters when the same application does both quick lookups and hard problems. tool-calling in openai's harmony format is strong.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free (apache 2.0)
website
github.com
our verdict

the most capable model here if you have the memory, with the genuinely useful ability to turn reasoning effort down when a task doesn't need it.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more small on-device llms

TinyLlama

1.1b model trained on 3t tokens, a common baseline for tiny deployments

#baseline#1b

MobileLLM

meta's research line for sub-billion-parameter on-device models

#research#sub-1b

H2O Danube

h2o.ai's small models released for edge and offline use

#edge

Gemma 4 (E2B / E4B / 12B)

we checked thisfree (apache 2.0)

text, images and audio on a phone, under an apache licence google took three generations to arrive at.

#anything-that-needs-to-s#on-hardware-you-already-