MiniMax M2.7 ranks #14 of 15 in our open-weight llms testing. announced as open source, then relicensed after release to forbid commercial use without written permission..
55/100
a capable agentic model you cannot ship on. the licence changed after the headlines were written, which is exactly why we read the file rather than the announcement.
why people look for an alternative
−commercial use needs minimax's prior written permission
−licence was revised after the 'open sources' announcement
−predecessors were mit, setting a false expectation
−mandatory attribution even once authorised
stay with MiniMax M2.7 if strong agentic engineering scores for its size is the thing you care about most — nothing below beats it on that.
#1 in open-weight llms · the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.
95/100
verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.
GLM-5.2 vs MiniMax M2.7
MiniMax M2.7
GLM-5.2
price
free for non-commercial only
free (mit)
free tier
yes
yes
license
non-commercial
mit
params
230b total / 10b active
753b total / ~40b active
context
~200k tokens
1m tokens
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
8x h200 node
switch foranyone who wants the strongest open model available and needs the legal answer to be boring.
pros
+plain unmodified mit — no traps of any kind
+top-ranked open-weight model on the artificial analysis index
+one-million-token context
+launch claims and licence file actually agree
cons
−roughly 750gb of gpu memory at fp8 — a full node
−active-parameter count only confirmed via secondary sources
#2 in open-weight llms · 1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.
92/100
verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.
DeepSeek V4 Pro vs MiniMax M2.7
MiniMax M2.7
DeepSeek V4 Pro
price
free for non-commercial only
free (mit)
free tier
yes
yes
license
non-commercial
mit
params
230b total / 10b active
1.6t total / 49b active
context
~200k tokens
1m tokens
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
multi-node cluster
switch forcoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.
pros
+plain mit at genuine frontier coding capability
+80.6% on swe-bench verified per vendor reporting
+one-million-token context
+cheapest frontier hosting of any model here
cons
−still labelled preview; promised stable release never shipped
−roughly 1.9tb at fp16 — multi-node self-hosting
−tool-use and multimodal breadth trail closed rivals
#3 in open-weight llms · europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence.
89/100
verdictthe highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.
Mistral Large 3 vs MiniMax M2.7
MiniMax M2.7
Mistral Large 3
price
free for non-commercial only
free (apache 2.0)
free tier
yes
yes
license
non-commercial
apache 2.0
params
230b total / 10b active
675b total / 41b active
context
~200k tokens
256k tokens
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
8x80gb cluster
switch forteams who want top-tier open coding performance from a european vendor with unambiguous terms.
pros
+apache 2.0 with no custom mistral terms attached
+top open-source coding model on lmarena
+full size ladder from 3b to 675b under one licence
+256k context
cons
−shipped december 2025 — the oldest of the top five here
−a larger successor is in early access but unreleased
−675b total puts self-hosting out of individual reach
#4 in open-weight llms · a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.
86/100
verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.
Kimi K2.7 Code vs MiniMax M2.7
MiniMax M2.7
Kimi K2.7 Code
price
free for non-commercial only
free (modified mit)
free tier
yes
yes
license
non-commercial
modified mit
params
230b total / 10b active
1t total / 32b active
context
~200k tokens
256k tokens
commercial use
needs written permission
attribution above scale
runs on
multi-gpu node
2x a100 at int4
switch forlong-running agentic engineering work, for anyone comfortably below the attribution thresholds.
pros
+attribution clause only triggers at 100m users or $20m monthly revenue
+purpose-built for long-horizon agentic engineering
+~30% fewer thinking tokens than its predecessor
+int4 deployment reportedly fits two a100 80gb cards
cons
−'modified mit' is described as plain mit almost everywhere
−trails frontier closed models on its vendor's own benchmarks
#5 in open-weight llms · thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.
84/100
verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.
Inkling vs MiniMax M2.7
MiniMax M2.7
Inkling
price
free for non-commercial only
free (apache 2.0)
free tier
yes
yes
license
non-commercial
apache 2.0
params
230b total / 10b active
975b total / 41b active
context
~200k tokens
1m tokens
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
8x80gb+ cluster
switch forteams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.
pros
+apache 2.0 from a major new us lab
+genuinely multimodal — text, image, audio and video training
+up to 1m context from downloaded weights
+unusually candid about its own limitations
cons
−licence confirmed from the model card, not the raw repo file
−behind deepseek v4 pro and kimi k2.6 on the intelligence index
−975b total — enterprise-scale hosting only
−tinker api caps context at 256k, lower than the weights allow
#8 in open-weight llms · the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.
79/100
verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.
Gemma 4 vs MiniMax M2.7
MiniMax M2.7
Gemma 4
price
free for non-commercial only
free (apache 2.0)
free tier
yes
yes
license
non-commercial
apache 2.0
params
230b total / 10b active
31b dense / 26b moe
context
~200k tokens
256k tokens
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
single consumer gpu
switch foron-device and laptop-class multimodal work where the model has to run locally and legally.
pros
+genuine apache 2.0, replacing three generations of custom terms
+multimodal variants that run on laptops and edge devices
+256k context and 140+ languages
+31b dense fits a single consumer gpu when quantised
cons
−repos still gated behind an accept-terms click-through
−google's old custom terms page is still live and unclarified
#9 in open-weight llms · genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand.
77/100
verdictan unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.
Qwen3.6-35B-A3B vs MiniMax M2.7
MiniMax M2.7
Qwen3.6-35B-A3B
price
free for non-commercial only
free (apache 2.0)
free tier
yes
yes
license
non-commercial
apache 2.0
params
230b total / 10b active
35b total / 3b active
context
~200k tokens
not confirmed
commercial use
needs written permission
unrestricted
runs on
multi-gpu node
single high-end gpu
switch foragentic coding on modest hardware, for teams who want alibaba's engineering without alibaba's api.
pros
+standard unmodified apache 2.0
+3b active — runs on one prosumer gpu
+vendor-reported 73.4 on swe-bench for its size
+combined thinking and non-thinking modes
cons
−constantly conflated with the closed api-only qwen3.7-max
−35b total means limited knowledge breadth
−context length not confirmed from a primary source
#10 in open-weight llms · a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.
74/100
verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.
NVIDIA Nemotron 3 Ultra vs MiniMax M2.7
MiniMax M2.7
NVIDIA Nemotron 3 Ultra
price
free for non-commercial only
free (licence disputed)
free tier
yes
yes
license
non-commercial
openmdw-1.1 or custom — disputed
params
230b total / 10b active
550b total / 55b active
context
~200k tokens
1m tokens
commercial use
needs written permission
unresolved
runs on
multi-gpu node
multi-node cluster
switch forenterprises with legal teams who can get nvidia to say which licence actually governs before deployment.
pros
+550b/55b active with a one-million-token context
+built for long-horizon agentic reasoning and tool-calling
+openmdw-1.1, if it governs, is genuinely permissive
cons
−nvidia's own pages name two different licences for it
−the custom licence adds a naming mandate and indemnification
every tool on this page went through the same test as MiniMax M2.7 — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.