Qwen3.6-35B-A3B ranks #9 of 15 in our open-weight llms testing. genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand..
77/100
an unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.
why people look for an alternative
−constantly conflated with the closed api-only qwen3.7-max
−35b total means limited knowledge breadth
−context length not confirmed from a primary source
−benchmark figures are vendor-reported only
stay with Qwen3.6-35B-A3B if standard unmodified apache 2.0 is the thing you care about most — nothing below beats it on that.
#1 in open-weight llms · the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.
95/100
verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.
GLM-5.2 vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
GLM-5.2
price
free (apache 2.0)
free (mit)
free tier
yes
yes
license
apache 2.0
mit
params
35b total / 3b active
753b total / ~40b active
context
not confirmed
1m tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
8x h200 node
switch foranyone who wants the strongest open model available and needs the legal answer to be boring.
pros
+plain unmodified mit — no traps of any kind
+top-ranked open-weight model on the artificial analysis index
+one-million-token context
+launch claims and licence file actually agree
cons
−roughly 750gb of gpu memory at fp8 — a full node
−active-parameter count only confirmed via secondary sources
#2 in open-weight llms · 1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.
92/100
verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.
DeepSeek V4 Pro vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
DeepSeek V4 Pro
price
free (apache 2.0)
free (mit)
free tier
yes
yes
license
apache 2.0
mit
params
35b total / 3b active
1.6t total / 49b active
context
not confirmed
1m tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
multi-node cluster
switch forcoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.
pros
+plain mit at genuine frontier coding capability
+80.6% on swe-bench verified per vendor reporting
+one-million-token context
+cheapest frontier hosting of any model here
cons
−still labelled preview; promised stable release never shipped
−roughly 1.9tb at fp16 — multi-node self-hosting
−tool-use and multimodal breadth trail closed rivals
#3 in open-weight llms · europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence.
89/100
verdictthe highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.
Mistral Large 3 vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
Mistral Large 3
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
35b total / 3b active
675b total / 41b active
context
not confirmed
256k tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
8x80gb cluster
switch forteams who want top-tier open coding performance from a european vendor with unambiguous terms.
pros
+apache 2.0 with no custom mistral terms attached
+top open-source coding model on lmarena
+full size ladder from 3b to 675b under one licence
+256k context
cons
−shipped december 2025 — the oldest of the top five here
−a larger successor is in early access but unreleased
−675b total puts self-hosting out of individual reach
#4 in open-weight llms · a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.
86/100
verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.
Kimi K2.7 Code vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
Kimi K2.7 Code
price
free (apache 2.0)
free (modified mit)
free tier
yes
yes
license
apache 2.0
modified mit
params
35b total / 3b active
1t total / 32b active
context
not confirmed
256k tokens
commercial use
unrestricted
attribution above scale
runs on
single high-end gpu
2x a100 at int4
switch forlong-running agentic engineering work, for anyone comfortably below the attribution thresholds.
pros
+attribution clause only triggers at 100m users or $20m monthly revenue
+purpose-built for long-horizon agentic engineering
+~30% fewer thinking tokens than its predecessor
+int4 deployment reportedly fits two a100 80gb cards
cons
−'modified mit' is described as plain mit almost everywhere
−trails frontier closed models on its vendor's own benchmarks
#5 in open-weight llms · thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.
84/100
verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.
Inkling vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
Inkling
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
35b total / 3b active
975b total / 41b active
context
not confirmed
1m tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
8x80gb+ cluster
switch forteams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.
pros
+apache 2.0 from a major new us lab
+genuinely multimodal — text, image, audio and video training
+up to 1m context from downloaded weights
+unusually candid about its own limitations
cons
−licence confirmed from the model card, not the raw repo file
−behind deepseek v4 pro and kimi k2.6 on the intelligence index
−975b total — enterprise-scale hosting only
−tinker api caps context at 256k, lower than the weights allow
#8 in open-weight llms · the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.
79/100
verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.
Gemma 4 vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
Gemma 4
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
35b total / 3b active
31b dense / 26b moe
context
not confirmed
256k tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
single consumer gpu
switch foron-device and laptop-class multimodal work where the model has to run locally and legally.
pros
+genuine apache 2.0, replacing three generations of custom terms
+multimodal variants that run on laptops and edge devices
+256k context and 140+ languages
+31b dense fits a single consumer gpu when quantised
cons
−repos still gated behind an accept-terms click-through
−google's old custom terms page is still live and unclarified
#10 in open-weight llms · a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.
74/100
verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.
NVIDIA Nemotron 3 Ultra vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
NVIDIA Nemotron 3 Ultra
price
free (apache 2.0)
free (licence disputed)
free tier
yes
yes
license
apache 2.0
openmdw-1.1 or custom — disputed
params
35b total / 3b active
550b total / 55b active
context
not confirmed
1m tokens
commercial use
unrestricted
unresolved
runs on
single high-end gpu
multi-node cluster
switch forenterprises with legal teams who can get nvidia to say which licence actually governs before deployment.
pros
+550b/55b active with a one-million-token context
+built for long-horizon agentic reasoning and tool-calling
+openmdw-1.1, if it governs, is genuinely permissive
cons
−nvidia's own pages name two different licences for it
−the custom licence adds a naming mandate and indemnification
#11 in open-weight llms · the only model here where 'fully open' survives contact with the evidence — weights, training data, code and checkpoints.
72/100
verdictevery other entry calls itself open and means the weights. ai2 publishes the training data too — which is worth more than the benchmark places it gives up.
OLMo 3 vs Qwen3.6-35B-A3B
Qwen3.6-35B-A3B
OLMo 3
price
free (apache 2.0)
free (apache 2.0)
free tier
yes
yes
license
apache 2.0
apache 2.0
params
35b total / 3b active
7b and 32b dense
context
not confirmed
32k tokens
commercial use
unrestricted
unrestricted
runs on
single high-end gpu
single consumer gpu
switch forresearch, auditing, and anyone who needs to know what a model was actually trained on.
pros
+training data, code and checkpoints published, not just weights
+clean apache 2.0
+7b variant runs on a consumer gpu
+think and rlzero variants built for reasoning research
cons
−32k context — the shortest in this ranking
−benchmark ceiling well below the frontier open models
every tool on this page went through the same test as Qwen3.6-35B-A3B — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.