Qwen3.5-4B-Instruct
262k of context in about three gigabytes, apache 2.0, and two generations newer than the qwen everyone links to.
the best reasoning-per-gigabyte here and the longest context by a wide margin — just make sure you're downloading the version alibaba actually ships now.
262,144 tokens of native context in a model that needs about three gigabytes at four-bit is the standout number in this category. most rivals sit at 128k or below, and the ones that match it are larger. alibaba reports a nine-point gain on its own intelligence index over qwen3-4b at identical size and licence — its figure, not an independent one.
the licence is standard apache 2.0 with no appendix, consistent across the qwen3 and qwen3.5 dense lines, which makes it one of the cleanest commercial stories here.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- small on-device llms
- pricing
- free (apache 2.0)
- website
- huggingface.co
the best reasoning-per-gigabyte here and the longest context by a wide margin — just make sure you're downloading the version alibaba actually ships now.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.