Kimi K3 ranks #15 of 15 in our open-weight llms testing. billed as the largest open-weight model ever built, with the weights due tomorrow and a countdown page where the files should be..
48/100
last place is about availability, not merit: as of today there is nothing to download and no terms to read, and this ranking puts the licence first for everything on it.
why people look for an alternative
−no weights published as of this review
−no licence file exists yet, so the terms are unknown
−no independent benchmark or leaderboard standing possible
−hardware requirements are press estimates only
stay with Kimi K3 if claimed largest open-weight model ever built at ~2.8t parameters is the thing you care about most — nothing below beats it on that.
#1 in open-weight llms · the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.
95/100
verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.
GLM-5.2 vs Kimi K3
Kimi K3
GLM-5.2
price
unreleased
free (mit)
free tier
no
yes
license
none published yet
mit
params
~2.8t total (moe)
753b total / ~40b active
context
1m tokens
1m tokens
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
8x h200 node
switch foranyone who wants the strongest open model available and needs the legal answer to be boring.
pros
+plain unmodified mit — no traps of any kind
+top-ranked open-weight model on the artificial analysis index
+one-million-token context
+launch claims and licence file actually agree
cons
−roughly 750gb of gpu memory at fp8 — a full node
−active-parameter count only confirmed via secondary sources
#2 in open-weight llms · 1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.
92/100
verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.
DeepSeek V4 Pro vs Kimi K3
Kimi K3
DeepSeek V4 Pro
price
unreleased
free (mit)
free tier
no
yes
license
none published yet
mit
params
~2.8t total (moe)
1.6t total / 49b active
context
1m tokens
1m tokens
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
multi-node cluster
switch forcoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.
pros
+plain mit at genuine frontier coding capability
+80.6% on swe-bench verified per vendor reporting
+one-million-token context
+cheapest frontier hosting of any model here
cons
−still labelled preview; promised stable release never shipped
−roughly 1.9tb at fp16 — multi-node self-hosting
−tool-use and multimodal breadth trail closed rivals
#3 in open-weight llms · europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence.
89/100
verdictthe highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.
Mistral Large 3 vs Kimi K3
Kimi K3
Mistral Large 3
price
unreleased
free (apache 2.0)
free tier
no
yes
license
none published yet
apache 2.0
params
~2.8t total (moe)
675b total / 41b active
context
1m tokens
256k tokens
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
8x80gb cluster
switch forteams who want top-tier open coding performance from a european vendor with unambiguous terms.
pros
+apache 2.0 with no custom mistral terms attached
+top open-source coding model on lmarena
+full size ladder from 3b to 675b under one licence
+256k context
cons
−shipped december 2025 — the oldest of the top five here
−a larger successor is in early access but unreleased
−675b total puts self-hosting out of individual reach
#4 in open-weight llms · a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.
86/100
verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.
Kimi K2.7 Code vs Kimi K3
Kimi K3
Kimi K2.7 Code
price
unreleased
free (modified mit)
free tier
no
yes
license
none published yet
modified mit
params
~2.8t total (moe)
1t total / 32b active
context
1m tokens
256k tokens
commercial use
no licence to read
attribution above scale
runs on
~1.4tb at 4-bit (estimated)
2x a100 at int4
switch forlong-running agentic engineering work, for anyone comfortably below the attribution thresholds.
pros
+attribution clause only triggers at 100m users or $20m monthly revenue
+purpose-built for long-horizon agentic engineering
+~30% fewer thinking tokens than its predecessor
+int4 deployment reportedly fits two a100 80gb cards
cons
−'modified mit' is described as plain mit almost everywhere
−trails frontier closed models on its vendor's own benchmarks
#5 in open-weight llms · thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.
84/100
verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.
Inkling vs Kimi K3
Kimi K3
Inkling
price
unreleased
free (apache 2.0)
free tier
no
yes
license
none published yet
apache 2.0
params
~2.8t total (moe)
975b total / 41b active
context
1m tokens
1m tokens
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
8x80gb+ cluster
switch forteams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.
pros
+apache 2.0 from a major new us lab
+genuinely multimodal — text, image, audio and video training
+up to 1m context from downloaded weights
+unusually candid about its own limitations
cons
−licence confirmed from the model card, not the raw repo file
−behind deepseek v4 pro and kimi k2.6 on the intelligence index
−975b total — enterprise-scale hosting only
−tinker api caps context at 256k, lower than the weights allow
#8 in open-weight llms · the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.
79/100
verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.
Gemma 4 vs Kimi K3
Kimi K3
Gemma 4
price
unreleased
free (apache 2.0)
free tier
no
yes
license
none published yet
apache 2.0
params
~2.8t total (moe)
31b dense / 26b moe
context
1m tokens
256k tokens
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
single consumer gpu
switch foron-device and laptop-class multimodal work where the model has to run locally and legally.
pros
+genuine apache 2.0, replacing three generations of custom terms
+multimodal variants that run on laptops and edge devices
+256k context and 140+ languages
+31b dense fits a single consumer gpu when quantised
cons
−repos still gated behind an accept-terms click-through
−google's old custom terms page is still live and unclarified
#9 in open-weight llms · genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand.
77/100
verdictan unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.
Qwen3.6-35B-A3B vs Kimi K3
Kimi K3
Qwen3.6-35B-A3B
price
unreleased
free (apache 2.0)
free tier
no
yes
license
none published yet
apache 2.0
params
~2.8t total (moe)
35b total / 3b active
context
1m tokens
not confirmed
commercial use
no licence to read
unrestricted
runs on
~1.4tb at 4-bit (estimated)
single high-end gpu
switch foragentic coding on modest hardware, for teams who want alibaba's engineering without alibaba's api.
pros
+standard unmodified apache 2.0
+3b active — runs on one prosumer gpu
+vendor-reported 73.4 on swe-bench for its size
+combined thinking and non-thinking modes
cons
−constantly conflated with the closed api-only qwen3.7-max
−35b total means limited knowledge breadth
−context length not confirmed from a primary source
#10 in open-weight llms · a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.
74/100
verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.
NVIDIA Nemotron 3 Ultra vs Kimi K3
Kimi K3
NVIDIA Nemotron 3 Ultra
price
unreleased
free (licence disputed)
free tier
no
yes
license
none published yet
openmdw-1.1 or custom — disputed
params
~2.8t total (moe)
550b total / 55b active
context
1m tokens
1m tokens
commercial use
no licence to read
unresolved
runs on
~1.4tb at 4-bit (estimated)
multi-node cluster
switch forenterprises with legal teams who can get nvidia to say which licence actually governs before deployment.
pros
+550b/55b active with a one-million-token context
+built for long-horizon agentic reasoning and tool-calling
+openmdw-1.1, if it governs, is genuinely permissive
cons
−nvidia's own pages name two different licences for it
−the custom licence adds a naming mandate and indemnification
every tool on this page went through the same test as Kimi K3 — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.