NVIDIA Nemotron 3 Nano 4B
runs without a gpu at all, under a licence nvidia can rewrite whenever it likes.
impressive engineering — a hybrid mamba architecture built for machines without gpus — attached to the most conditional licence in this ranking.
the technical story is genuinely interesting. it's a hybrid mamba-2 design with just four attention layers, compressed from the 9b v2 model through nvidia's nemotron elastic framework, and explicitly built for cpu-only inference. 262k of context at around 2.5gb without needing a graphics card is a combination nothing else here offers.
the licence has teeth that apache 2.0 doesn't. commercial use and redistribution are permitted and nvidia disclaims ownership of outputs — but your rights terminate if you sue over the model, they terminate if you 'bypass, disable, reduce the efficacy of, or circumvent' any safety guardrail in it, and nvidia reserves the right to update the licence unilaterally to meet legal or regulatory requirements, leaving you to accept the new terms or stop using and distributing it.
the description above is ours, condensed from the ranking. pricing moves — check it on the vendor's own page before you rely on it.
- category
- small on-device llms
- pricing
- free (nemotron open model license)
- website
- nvidia.com
impressive engineering — a hybrid mamba architecture built for machines without gpus — attached to the most conditional licence in this ranking.
we researched this category against vendors' own pricing pages and licence files. that is where this line comes from — not from the vendor, and not from anything they paid for.