gpt-oss-120b ranks #6 of 15 in our open-weight llms testing. the only genuinely capable model here that fits on a single 80gb card, under clean apache 2.0..
82/100
nearly a year old and still the most practical self-host on this list — one card, no licence questions, real agentic ability.
why people look for an alternative
−text-only in a multimodal field
−128k context, well below the 2026 frontier
−no successor since august 2025
−passed on the intelligence index by newer open releases
stay with gpt-oss-120b if unmodified apache 2.0 with no layered openai terms is the thing you care about most — nothing below beats it on that.
#1 in open-weight llms · the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.
95/100
verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.
GLM-5.2 vs gpt-oss-120b
gpt-oss-120b
GLM-5.2
price
free (apache 2.0)
free (mit)
free tier
yes
yes
license
apache 2.0
mit
params
117b total / 5.1b active
753b total / ~40b active
context
128k tokens
1m tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
8x h200 node
switch foranyone who wants the strongest open model available and needs the legal answer to be boring.
pros
+plain unmodified mit — no traps of any kind
+top-ranked open-weight model on the artificial analysis index
+one-million-token context
+launch claims and licence file actually agree
cons
−roughly 750gb of gpu memory at fp8 — a full node
−active-parameter count only confirmed via secondary sources
#2 in open-weight llms · 1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.
92/100
verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.
DeepSeek V4 Pro vs gpt-oss-120b
gpt-oss-120b
DeepSeek V4 Pro
price
free (apache 2.0)
free (mit)
free tier
yes
yes
license
apache 2.0
mit
params
117b total / 5.1b active
1.6t total / 49b active
context
128k tokens
1m tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
multi-node cluster
switch forcoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.
pros
+plain mit at genuine frontier coding capability
+80.6% on swe-bench verified per vendor reporting
+one-million-token context
+cheapest frontier hosting of any model here
cons
−still labelled preview; promised stable release never shipped
−roughly 1.9tb at fp16 — multi-node self-hosting
−tool-use and multimodal breadth trail closed rivals
#3 in open-weight llms · europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence.
89/100
verdictthe highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.
Mistral Large 3 vs gpt-oss-120b
gpt-oss-120b
Mistral Large 3
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
117b total / 5.1b active
675b total / 41b active
context
128k tokens
256k tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
8x80gb cluster
switch forteams who want top-tier open coding performance from a european vendor with unambiguous terms.
pros
+apache 2.0 with no custom mistral terms attached
+top open-source coding model on lmarena
+full size ladder from 3b to 675b under one licence
+256k context
cons
−shipped december 2025 — the oldest of the top five here
−a larger successor is in early access but unreleased
−675b total puts self-hosting out of individual reach
#4 in open-weight llms · a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.
86/100
verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.
Kimi K2.7 Code vs gpt-oss-120b
gpt-oss-120b
Kimi K2.7 Code
price
free (apache 2.0)
free (modified mit)
free tier
yes
yes
license
apache 2.0
modified mit
params
117b total / 5.1b active
1t total / 32b active
context
128k tokens
256k tokens
commercial use
unrestricted
attribution above scale
runs on
single 80gb gpu
2x a100 at int4
switch forlong-running agentic engineering work, for anyone comfortably below the attribution thresholds.
pros
+attribution clause only triggers at 100m users or $20m monthly revenue
+purpose-built for long-horizon agentic engineering
+~30% fewer thinking tokens than its predecessor
+int4 deployment reportedly fits two a100 80gb cards
cons
−'modified mit' is described as plain mit almost everywhere
−trails frontier closed models on its vendor's own benchmarks
#5 in open-weight llms · thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.
84/100
verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.
Inkling vs gpt-oss-120b
gpt-oss-120b
Inkling
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
117b total / 5.1b active
975b total / 41b active
context
128k tokens
1m tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
8x80gb+ cluster
switch forteams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.
pros
+apache 2.0 from a major new us lab
+genuinely multimodal — text, image, audio and video training
+up to 1m context from downloaded weights
+unusually candid about its own limitations
cons
−licence confirmed from the model card, not the raw repo file
−behind deepseek v4 pro and kimi k2.6 on the intelligence index
−975b total — enterprise-scale hosting only
−tinker api caps context at 256k, lower than the weights allow
#8 in open-weight llms · the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.
79/100
verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.
Gemma 4 vs gpt-oss-120b
gpt-oss-120b
Gemma 4
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
117b total / 5.1b active
31b dense / 26b moe
context
128k tokens
256k tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
single consumer gpu
switch foron-device and laptop-class multimodal work where the model has to run locally and legally.
pros
+genuine apache 2.0, replacing three generations of custom terms
+multimodal variants that run on laptops and edge devices
+256k context and 140+ languages
+31b dense fits a single consumer gpu when quantised
cons
−repos still gated behind an accept-terms click-through
−google's old custom terms page is still live and unclarified
#9 in open-weight llms · genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand.
77/100
verdictan unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.
Qwen3.6-35B-A3B vs gpt-oss-120b
gpt-oss-120b
Qwen3.6-35B-A3B
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
117b total / 5.1b active
35b total / 3b active
context
128k tokens
not confirmed
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
single high-end gpu
switch foragentic coding on modest hardware, for teams who want alibaba's engineering without alibaba's api.
pros
+standard unmodified apache 2.0
+3b active — runs on one prosumer gpu
+vendor-reported 73.4 on swe-bench for its size
+combined thinking and non-thinking modes
cons
−constantly conflated with the closed api-only qwen3.7-max
−35b total means limited knowledge breadth
−context length not confirmed from a primary source
#10 in open-weight llms · a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.
74/100
verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.
NVIDIA Nemotron 3 Ultra vs gpt-oss-120b
gpt-oss-120b
NVIDIA Nemotron 3 Ultra
price
free (apache 2.0)
free (licence disputed)
free tier
yes
yes
license
apache 2.0
openmdw-1.1 or custom — disputed
params
117b total / 5.1b active
550b total / 55b active
context
128k tokens
1m tokens
commercial use
unrestricted
unresolved
runs on
single 80gb gpu
multi-node cluster
switch forenterprises with legal teams who can get nvidia to say which licence actually governs before deployment.
pros
+550b/55b active with a one-million-token context
+built for long-horizon agentic reasoning and tool-calling
+openmdw-1.1, if it governs, is genuinely permissive
cons
−nvidia's own pages name two different licences for it
−the custom licence adds a naming mandate and indemnification
#11 in open-weight llms · the only model here where 'fully open' survives contact with the evidence — weights, training data, code and checkpoints.
72/100
verdictevery other entry calls itself open and means the weights. ai2 publishes the training data too — which is worth more than the benchmark places it gives up.
OLMo 3 vs gpt-oss-120b
gpt-oss-120b
OLMo 3
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
117b total / 5.1b active
7b and 32b dense
context
128k tokens
32k tokens
commercial use
unrestricted
unrestricted
runs on
single 80gb gpu
single consumer gpu
switch forresearch, auditing, and anyone who needs to know what a model was actually trained on.
pros
+training data, code and checkpoints published, not just weights
+clean apache 2.0
+7b variant runs on a consumer gpu
+think and rlzero variants built for reasoning research
cons
−32k context — the shortest in this ranking
−benchmark ceiling well below the frontier open models
every tool on this page went through the same test as gpt-oss-120b — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.