verifier.org

Phi-4-mini-instruct

plain mit, strong function calling, and the oldest model in this ranking by a year.

free (mit)#structured-output#tool-calling-in-a-small-#where-reliability-beats-

clean mit terms and genuinely good instruction following for 3.8b — held back by being a february 2025 model in a field that has moved twice since.

the licence is plain mit with no field-of-use limits, revenue cap or naming mandate, which is the simplest commercial story on this page. at 3.8b and roughly 2.5gb quantised it runs comfortably on any modern laptop, with 128k context and an unusually large 200k-token vocabulary that helps on multilingual input.

what it does well is follow instructions and call functions reliably at a size where that usually degrades. for pipelines that need structured output rather than conversation, that reliability is worth more than a benchmark point.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free (mit)
our verdict

clean mit terms and genuinely good instruction following for 3.8b — held back by being a february 2025 model in a field that has moved twice since.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more small on-device llms

TinyLlama

1.1b model trained on 3t tokens, a common baseline for tiny deployments

#baseline#1b

MobileLLM

meta's research line for sub-billion-parameter on-device models

#research#sub-1b

H2O Danube

h2o.ai's small models released for edge and offline use

#edge

Gemma 4 (E2B / E4B / 12B)

we checked thisfree (apache 2.0)

text, images and audio on a phone, under an apache licence google took three generations to arrive at.

#anything-that-needs-to-s#on-hardware-you-already-