verifier.org

Qwen3.5-4B-Instruct

262k of context in about three gigabytes, apache 2.0, and two generations newer than the qwen everyone links to.

free (apache 2.0)#long-document-work-on-a-#where-the-context-window

the best reasoning-per-gigabyte here and the longest context by a wide margin — just make sure you're downloading the version alibaba actually ships now.

262,144 tokens of native context in a model that needs about three gigabytes at four-bit is the standout number in this category. most rivals sit at 128k or below, and the ones that match it are larger. alibaba reports a nine-point gain on its own intelligence index over qwen3-4b at identical size and licence — its figure, not an independent one.

the licence is standard apache 2.0 with no appendix, consistent across the qwen3 and qwen3.5 dense lines, which makes it one of the cleanest commercial stories here.

the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.

pricing
free (apache 2.0)
our verdict

the best reasoning-per-gigabyte here and the longest context by a wide margin — just make sure you're downloading the version alibaba actually ships now.

we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.

more small on-device llms

TinyLlama

1.1b model trained on 3t tokens, a common baseline for tiny deployments

#baseline#1b

MobileLLM

meta's research line for sub-billion-parameter on-device models

#research#sub-1b

H2O Danube

h2o.ai's small models released for edge and offline use

#edge

Gemma 4 (E2B / E4B / 12B)

we checked thisfree (apache 2.0)

text, images and audio on a phone, under an apache licence google took three generations to arrive at.

#anything-that-needs-to-s#on-hardware-you-already-