verifier.org

best 15 open-weight llms

ranked on what the licence file actually says, then on capability and what it takes to run — because in 2026 'open source' in a launch post and the terms in the repo are routinely different documents.

last reviewed 26 jul 2026 · 15 tools tested ·list curated by Onur Ozcanxin

the short version
best overallGLM-5.2anyone who wants the strongest open model available and needs the legal answer to be boring.95/100runner-upDeepSeek V4 Procoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.92/100

we read the licence file for every model here rather than the announcement, and the gap between the two is the story of this category. minimax announced 'minimax open sources m2.7' in april, its two previous models shipped under plain mit, and the licence in the repo was then revised to prohibit commercial use without minimax's prior written authorisation. teams who benchmarked it on the headline are shipping into breach.

the traps are rarely outright bans. kimi k2.7 code is 'modified mit', and the modification is a clause requiring you to display 'kimi k2.7 code' in your product's interface once you pass 100 million monthly users or $20 million in monthly revenue. llama 4 carries a 700-million-user gate, a mandate that anything you derive from it be named starting with 'llama', and a clause terminating your rights if you bring ip litigation against meta. none of that appears in a launch blog.

the good news is that the top of this list is genuinely clean. glm-5.2 and deepseek v4 are plain, unmodified mit. mistral large 3, gpt-oss-120b, olmo 3, inkling and — this is the year's biggest licensing improvement — gemma 4 are apache 2.0. google's gemma shipped under custom terms with a prohibited-use policy for three generations and quietly became genuinely open with this one.

one model is deliberately absent. qwen3.7-max has no weights, no hugging face repository and no committed release date, which does not stop roundups listing it beside the genuinely apache-2.0 qwen3.6 line — those are two different products and only one of them is open. we'll rank it when you can download it.

kimi k3 is included and ranked last, which needs explaining. it was announced on 16 july as the largest open-weight model ever built, its weights are due tomorrow, and its hugging face page is still a countdown placeholder. last place here reflects what you can do with it today — nothing — and not a judgement about the model, which may well deserve the top half of this page next week. we rank licence above benchmark throughout, and a model with no licence file yet cannot outrank one with clean mit.

advertisement
  1. 1

    GLM-5.2

    the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.

    95/100

    verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.

    best for
    anyone who wants the strongest open model available and needs the legal answer to be boring.
    price
    free (mit)
    pricing note
    weights downloadable; hosted at $1.40/$4.40 per million at four inference providers
    free tier
    yes
    license
    mit
    params
    753b total / ~40b active
    context
    1m tokens
    commercial use
    unrestricted
    runs on
    8x h200 node

    zhipu's launch material claims 'pure open' with 'no regional limits', and the licence file is standard unmodified mit text with a 2026 copyright line. no revenue cap, no field-of-use clause, no geographic carve-out, no attribution mandate beyond keeping the notice. of the fourteen models here, this is one of the few where the marketing and the legal text say the same thing.

    it also currently sits at the top of the artificial analysis intelligence index among open-weight models, having displaced kimi k2.6. it is a 753b sparse mixture-of-experts with roughly 40b active, a full one-million-token context, and it is built for long-horizon agentic and coding work rather than chat.

    running it yourself is a datacentre exercise: community estimates put fp8 at around 750gb of gpu memory — an 8x h200 node — and full fp16 near 1.6tb, with int4 bringing it to roughly 400gb. most teams will rent it, where the price is identical at fireworks, baseten, novita and together.

    pros
    • +plain unmodified mit — no traps of any kind
    • +top-ranked open-weight model on the artificial analysis index
    • +one-million-token context
    • +launch claims and licence file actually agree
    cons
    • roughly 750gb of gpu memory at fp8 — a full node
    • active-parameter count only confirmed via secondary sources
    • no published vendor statement of its weaknesses
  2. 2

    DeepSeek V4 Pro

    1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.

    92/100

    verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.

    best for
    coding and reasoning workloads where you want frontier results without a frontier bill or a licence review.
    price
    free (mit)
    pricing note
    weights downloadable; hosted from $1.30/$2.60 per million at deepinfra
    free tier
    yes
    license
    mit
    params
    1.6t total / 49b active
    context
    1m tokens
    commercial use
    unrestricted
    runs on
    multi-node cluster

    the licence is plain mit, fetched from the repository rather than inferred: no revenue cap, no field-of-use restriction, no attribution requirement beyond the copyright notice. for a model reported at 80.6% on swe-bench verified and ahead of gpt-5.5 on codeforces and livecodebench, that is a remarkable thing to be able to write.

    the architecture is a 1.6-trillion-parameter mixture-of-experts with 49b active and a one-million-token context. where it lags is breadth rather than depth — tool-use ecosystem, multimodality and very long agentic loops still favour the closed frontier models.

    the caveat worth pausing on: this has been labelled a preview since april, and the stable build deepseek targeted for mid-july has not appeared in its changelog. that is a real consideration for anything going into production on downloaded weights, and it is why this sits second rather than first.

    pros
    • +plain mit at genuine frontier coding capability
    • +80.6% on swe-bench verified per vendor reporting
    • +one-million-token context
    • +cheapest frontier hosting of any model here
    cons
    • still labelled preview; promised stable release never shipped
    • roughly 1.9tb at fp16 — multi-node self-hosting
    • tool-use and multimodal breadth trail closed rivals
  3. 3

    Mistral Large 3

    europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence.

    89/100

    verdictthe highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.

    best for
    teams who want top-tier open coding performance from a european vendor with unambiguous terms.
    price
    free (apache 2.0)
    pricing note
    weights downloadable; smaller ministral 3 siblings at 14b, 8b and 3b are also apache 2.0
    free tier
    yes
    license
    apache 2.0
    params
    675b total / 41b active
    context
    256k tokens
    commercial use
    unrestricted
    runs on
    8x80gb cluster

    mistral has shipped flagships under its own research-and-commercial split licence before, so apache 2.0 here was not a foregone conclusion. the model card links the standard apache text directly and adds no custom terms — the announcement's 'all models released under apache 2.0' holds up against the repository.

    at 675b total and 41b active with a 256k context, it is reported as the top open-source coding model on lmarena and second overall among open models at roughly 1418 elo, behind only gemini among all models at publication. the ministral 3 siblings at 14b, 8b and 3b carry the same licence, which makes this the most complete permissive family here.

    two things temper it. december 2025 is a long time ago in this field, and mistral's ceo confirmed a larger sparse successor in early access in july that has not shipped. and 675b total means full-precision self-hosting is an organisational decision, not an individual one.

    pros
    • +apache 2.0 with no custom mistral terms attached
    • +top open-source coding model on lmarena
    • +full size ladder from 3b to 675b under one licence
    • +256k context
    cons
    • shipped december 2025 — the oldest of the top five here
    • a larger successor is in early access but unreleased
    • 675b total puts self-hosting out of individual reach
    advertisement
  4. 4

    Kimi K2.7 Code

    a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.

    86/100

    verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.

    best for
    long-running agentic engineering work, for anyone comfortably below the attribution thresholds.
    price
    free (modified mit)
    pricing note
    weights downloadable; attribution obligation triggers above 100m monthly users or $20m monthly revenue
    free tier
    yes
    license
    modified mit
    params
    1t total / 32b active
    context
    256k tokens
    commercial use
    attribution above scale
    runs on
    2x a100 at int4

    the modification to mit is a single scale-triggered clause: if your product passes 100 million monthly active users or $20 million in monthly revenue, you must display 'kimi k2.7 code' prominently in its interface. below that it behaves as plain mit — no revenue share, no field-of-use ban, no geographic limit. it is one of the milder terms in this category, and it is buried in the licence file rather than in any announcement.

    the model is a 1t-parameter mixture-of-experts with 32b active across 384 experts and a 256k context, tuned for long-horizon real-world software engineering. moonshot reports it using roughly 30% fewer thinking tokens than k2.6 for comparable completion, which matters when you're paying per token.

    it trails the closed frontier on raw scores — 62.0 against gpt-5.5's 69.0 on kimi's own code benchmark, 35.1 against claude opus's 42.8 on ml-research tasks — and sits at position 31 on the artificial analysis index. native int4 quantisation reportedly brings deployment down to around two a100 80gb cards, which is the most approachable footprint of any trillion-parameter model here.

    pros
    • +attribution clause only triggers at 100m users or $20m monthly revenue
    • +purpose-built for long-horizon agentic engineering
    • +~30% fewer thinking tokens than its predecessor
    • +int4 deployment reportedly fits two a100 80gb cards
    cons
    • 'modified mit' is described as plain mit almost everywhere
    • trails frontier closed models on its vendor's own benchmarks
    • position 31 on the artificial analysis index
    • 256k context, against 1m for the top three
  5. 5

    Inkling

    thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.

    84/100

    verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.

    best for
    teams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.
    price
    free (apache 2.0)
    pricing note
    weights downloadable at up to 1m context; 256k via the vendor's tinker api
    free tier
    yes
    license
    apache 2.0
    params
    975b total / 41b active
    context
    1m tokens
    commercial use
    unrestricted
    runs on
    8x80gb+ cluster

    released on 15 july, this is the first from-scratch open-weight model from mira murati's thinking machines lab: 975b total with 41b active, trained on 45 trillion tokens spanning text, image, audio and video, with up to a million tokens of context from the downloaded weights.

    it is explicitly designed as a customisation base rather than a finished product, paired with the vendor's tinker fine-tuning platform, and its stated design goals include resistance to censorship. the model card's own admission that it is 'not the strongest overall model available today, open or closed' — swe-bench verified at 77.6% against roughly 95% for leading closed models — is the kind of candour this category could use more of.

    one verification caveat we'd rather state than hide: apache 2.0 is confirmed from thinking machines' own model card, not from a raw licence file pulled out of the repository, which fell outside our budget. treat it as provisionally confirmed and check the repo before committing.

    pros
    • +apache 2.0 from a major new us lab
    • +genuinely multimodal — text, image, audio and video training
    • +up to 1m context from downloaded weights
    • +unusually candid about its own limitations
    cons
    • licence confirmed from the model card, not the raw repo file
    • behind deepseek v4 pro and kimi k2.6 on the intelligence index
    • 975b total — enterprise-scale hosting only
    • tinker api caps context at 256k, lower than the weights allow
  6. 6

    gpt-oss-120b

    the only genuinely capable model here that fits on a single 80gb card, under clean apache 2.0.

    82/100

    verdictnearly a year old and still the most practical self-host on this list — one card, no licence questions, real agentic ability.

    best for
    teams who want to actually own the deployment rather than rent it, on hardware they can afford.
    price
    free (apache 2.0)
    pricing note
    weights downloadable; ~60gb at native mxfp4 quantisation
    free tier
    yes
    license
    apache 2.0
    params
    117b total / 5.1b active
    context
    128k tokens
    commercial use
    unrestricted
    runs on
    single 80gb gpu

    openai's open-weight release carries unmodified apache 2.0. the only restrictions are apache's own patent-litigation termination clause and ordinary trademark limits; a separate usage policy exists but is explicitly not part of the licence and imposes nothing contractual. after the licences elsewhere in this category, that plainness is worth something.

    the practical case is the footprint. 117b total with only 5.1b active per token means it fits a single 80gb gpu at native mxfp4 — roughly 60gb of weights with room left for kv cache — and the small active path makes cpu and consumer-gpu inference viable at lower throughput. nothing else here with real agentic and tool-calling ability runs on one card.

    the age shows. released august 2025 with no successor, it is text-only while the 2026 field has gone multimodal, and its 128k context is a quarter of what the models above it offer. it has been passed on the intelligence index by inkling, gemma 4 and deepseek v4.

    pros
    • +unmodified apache 2.0 with no layered openai terms
    • +runs on a single 80gb gpu at mxfp4
    • +5.1b active path makes cpu inference viable
    • +strong agentic tool-calling with configurable reasoning effort
    cons
    • text-only in a multimodal field
    • 128k context, well below the 2026 frontier
    • no successor since august 2025
    • passed on the intelligence index by newer open releases
  7. 7

    DeepSeek V4 Flash

    a separate model rather than a trim of v4 pro — a fifth the size, same mit licence, a twelfth the hosted price.

    81/100

    verdictthe best price-to-capability ratio in the category, and widely mistaken for a variant of v4 pro when it's its own release with its own weights.

    best for
    high-volume work that doesn't need pro-tier reasoning but does need a licence you can ignore.
    price
    free (mit)
    pricing note
    weights downloadable; hosted at $0.14/$0.28 per million at fireworks — the cheapest frontier-adjacent rate we found
    free tier
    yes
    license
    mit
    params
    284b total / 13b active
    context
    1m tokens
    commercial use
    unrestricted
    runs on
    single 8x80gb node

    worth stating clearly because comparison tables get it wrong: v4 flash is a separate repository and a separate release at 284b total with 13b active, not a quantisation or distillation of v4 pro's 1.6t. it carries the same plain mit licence.

    the economics are the story. hosted at fireworks it runs $0.14 in and $0.28 out per million tokens against v4 pro's $1.74 and $3.48 — a twelfth the cost — and it scores 40.3% on the artificial analysis intelligence index, which for a model of this size and price is a strong showing.

    self-hosting is meaningfully more achievable than the pro tier too, with 13b active rather than 49b. if your workload is high-volume and doesn't need the last increment of reasoning, this is the entry on this page most likely to save you real money.

    pros
    • +plain mit, same as v4 pro
    • +cheapest frontier-adjacent hosted rate in the market
    • +13b active — far more self-hostable than the pro tier
    • +a fifth the total size for a fraction of the capability gap
    cons
    • routinely misdescribed as a variant of v4 pro
    • 40.3% on the intelligence index — clearly below the top tier
    • inherits v4's preview-status uncertainty
  8. 8

    Gemma 4

    the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.

    79/100

    verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.

    best for
    on-device and laptop-class multimodal work where the model has to run locally and legally.
    price
    free (apache 2.0)
    pricing note
    weights downloadable behind a gated click-through; family spans e2b and e4b through 12b, a 26b mixture-of-experts and a 31b dense variant
    free tier
    yes
    license
    apache 2.0
    params
    31b dense / 26b moe
    context
    256k tokens
    commercial use
    unrestricted
    runs on
    single consumer gpu

    gemma 1 through 3 shipped under google's own gemma terms of use: a prohibited-use policy, redistribution and notice obligations, downstream pass-through requirements and termination on breach. widely called 'open', never actually open source. gemma 4's hugging face card carries a plain apache-2.0 tag, and google's announcement says the same. that is a genuine and underreported change.

    the family is designed for local deployment: e2b and e4b variants handle text, image and audio on edge hardware, the 12b targets laptops, and the 31b dense and 26b sparse variants sit comfortably on one 24-48gb consumer card when quantised. 256k context and 140+ languages.

    two loose ends we could not close. the repositories remain gated behind an accept-terms click-through, and google's old gemma terms page is still live — we could not confirm whether that click-through still invokes it or whether the page is simply stale. and the training data stops at january 2025, which shows on anything recent.

    pros
    • +genuine apache 2.0, replacing three generations of custom terms
    • +multimodal variants that run on laptops and edge devices
    • +256k context and 140+ languages
    • +31b dense fits a single consumer gpu when quantised
    cons
    • repos still gated behind an accept-terms click-through
    • google's old custom terms page is still live and unclarified
    • training data cutoff of january 2025
  9. 9

    Qwen3.6-35B-A3B

    genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand.

    77/100

    verdictan unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.

    best for
    agentic coding on modest hardware, for teams who want alibaba's engineering without alibaba's api.
    price
    free (apache 2.0)
    pricing note
    weights downloadable; note that qwen3.7-max, the flagship tier, has no weights and no announced plan to release any
    free tier
    yes
    license
    apache 2.0
    params
    35b total / 3b active
    context
    not confirmed
    commercial use
    unrestricted
    runs on
    single high-end gpu

    the licence file is standard unmodified apache 2.0 with a 2026 alibaba cloud copyright and no appendix. at 35b total with just 3b active, alibaba describes it as performing on par with models ten times its active size, reporting 73.4 on swe-bench and 51.5 on terminal-bench 2.0 — vendor figures, not independently verified.

    the confusion worth avoiding: qwen3.7-max, announced in may as the flagship, has no hugging face repository, no weights and no committed release date. it is api-only through alibaba cloud and resold through openrouter and together. roundups that place '3.7' alongside the open qwen line are describing two different products, and only one of them is open.

    the honest limit is capacity. 35b total is a fraction of the frontier models above, so raw knowledge breadth and long-tail reasoning trail them. what you get instead is something that runs on a single high-end card.

    pros
    • +standard unmodified apache 2.0
    • +3b active — runs on one prosumer gpu
    • +vendor-reported 73.4 on swe-bench for its size
    • +combined thinking and non-thinking modes
    cons
    • constantly conflated with the closed api-only qwen3.7-max
    • 35b total means limited knowledge breadth
    • context length not confirmed from a primary source
    • benchmark figures are vendor-reported only
  10. 10

    NVIDIA Nemotron 3 Ultra

    a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.

    74/100

    verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.

    best for
    enterprises with legal teams who can get nvidia to say which licence actually governs before deployment.
    price
    free (licence disputed)
    pricing note
    hugging face links openmdw-1.1; nvidia's legal pages describe a custom nemotron open model licence with different obligations
    free tier
    yes
    license
    openmdw-1.1 or custom — disputed
    params
    550b total / 55b active
    context
    1m tokens
    commercial use
    unresolved
    runs on
    multi-node cluster

    the licence file linked from the hugging face card is the linux foundation's openmdw-1.1: no revenue cap, no field-of-use limit, no naming mandate, with an ip-litigation termination clause and responsibility on you to clear third-party rights. read alone, that is reasonably permissive.

    but nvidia's own legal-agreements pages describe nemotron models as governed by a custom nvidia nemotron open model licence, effective december 2025, which adds a naming mandate on derivatives and a broad user-indemnification clause that openmdw contains nothing like. we could not determine whether this model was quietly relicensed or the card is mislabelled, and nvidia's materials are inconsistent on it.

    the model itself is strong — 550b total with 55b active and a one-million-token context, built for long-context agentic reasoning and reliable tool-calling. it needs a multi-node cluster to serve, so the audience is exactly the kind of organisation that will notice an unresolved licence.

    pros
    • +550b/55b active with a one-million-token context
    • +built for long-horizon agentic reasoning and tool-calling
    • +openmdw-1.1, if it governs, is genuinely permissive
    cons
    • nvidia's own pages name two different licences for it
    • the custom licence adds a naming mandate and indemnification
    • 550b total — enterprise-only self-hosting
    • no verified public leaderboard standing found
  11. 11

    OLMo 3

    the only model here where 'fully open' survives contact with the evidence — weights, training data, code and checkpoints.

    72/100

    verdictevery other entry calls itself open and means the weights. ai2 publishes the training data too — which is worth more than the benchmark places it gives up.

    best for
    research, auditing, and anyone who needs to know what a model was actually trained on.
    price
    free (apache 2.0)
    pricing note
    weights, dolma 3 training corpus, dolci post-training data, training code and intermediate checkpoints all published
    free tier
    yes
    license
    apache 2.0
    params
    7b and 32b dense
    context
    32k tokens
    commercial use
    unrestricted
    runs on
    single consumer gpu

    apache 2.0 on the weights, and alongside them the dolma 3 training corpus, the dolci post-training datasets, the training and evaluation code, and the intermediate checkpoints — all linked from the model card. if you need to audit what went into a model, or reproduce it, this is the only real option on this page.

    the family is 7b and 32b dense, each in base, instruct, think and rlzero variants, which makes the reasoning-focused ones genuinely useful for research into how these models learn. the 7b runs on a 16-24gb consumer card and the 32b on a single 48-80gb card.

    the trade is capability. a 32b dense model does not compete with 675b and 1.6t sparse rivals, the 32k context is the shortest here, and it did not feature in july's intelligence index coverage. that is the honest cost of full transparency at a non-profit's budget.

    pros
    • +training data, code and checkpoints published, not just weights
    • +clean apache 2.0
    • +7b variant runs on a consumer gpu
    • +think and rlzero variants built for reasoning research
    cons
    • 32k context — the shortest in this ranking
    • benchmark ceiling well below the frontier open models
    • no current leaderboard standing found
    • released late 2025 with no 2026 update
  12. 12

    Phi-4 (reasoning-vision 15B)

    mit-licensed reasoning small enough for a laptop, with a context window that belongs to 2024.

    68/100

    verdictclean mit across the whole family and genuinely runnable anywhere — held back by a 16k context that rules out most 2026 workloads.

    best for
    edge and on-device reasoning where the model has to fit and the task is bounded.
    price
    free (mit)
    pricing note
    weights downloadable; family runs from 3.8b mini to this 15b reasoning-vision variant
    free tier
    yes
    license
    mit
    params
    15b dense
    context
    16k tokens
    commercial use
    unrestricted
    runs on
    single consumer gpu

    microsoft has kept every phi-4 variant under plain mit, from the 3.8b mini through the 14b reasoning models to this 15b reasoning-vision release from march. no field-of-use terms, no attribution beyond the copyright line. consistency across a seven-model family is rarer than it should be.

    the appeal is deployment reach: the 15b fits a single 24gb consumer card in bf16 or int8, and the mini variants run on cpu or entry-level edge hardware through ollama or foundry local. the reasoning was distilled from o3-mini traces and now extends to vision-grounded tasks.

    the 16,384-token context is the problem. against 256k and 1m elsewhere on this page, that rules out long documents, large codebases and most agentic loops. this is a model for bounded, well-specified tasks close to the device, and it is good at those.

    pros
    • +plain mit across all seven family variants
    • +15b fits a single 24gb consumer gpu
    • +3.8b mini runs on cpu and edge hardware
    • +vision-grounded reasoning at small scale
    cons
    • 16k context — by far the shortest here
    • narrow world knowledge versus larger 2026 models
    • not competitive on frontier benchmarks
    • licence read from the card's tag rather than a repo file
  13. 13

    Llama 4 Maverick

    the model that taught everyone to read licences, still carrying a user cap, a naming mandate and a litigation kill-switch.

    62/100

    verdictmeta calls it open source and it is not, by the definition's own terms — and after fifteen months it has been overtaken by models with cleaner licences and better scores.

    best for
    teams already standardised on llama tooling who have read the agreement and stay under the threshold.
    price
    free (llama 4 community licence)
    pricing note
    not an osi-approved licence; separate commercial licence required above 700m monthly active users
    free tier
    yes
    license
    llama 4 community licence
    params
    400b total / 17b active
    context
    1m tokens
    commercial use
    gated above 700m users
    runs on
    single 8x80gb node

    the llama 4 community licence carries four things worth knowing. any product above 700 million monthly active users needs a separate licence from meta. you must display 'built with llama', and any derivative model's name must begin with 'llama'. meta's acceptable use policy is incorporated by reference. and your rights terminate automatically if you bring ip litigation against meta over the model or its outputs.

    the benchmark history compounds it. meta submitted a differently-tuned, unreleased experimental checkpoint to lmarena in april 2025 that scored far above the weights it actually published — yann lecun acknowledged in january 2026 that 'results were fudged a little bit'. trackers now disagree wildly, with one aggregator placing it 205th of 215.

    the architecture — 400b total, 17b active, 1m context, multimodal — was competitive when it landed in april 2025. it now clearly trails deepseek v4 and qwen 3.6 on coding, and everything above it here has a cleaner licence.

    pros
    • +1m context and multimodal input
    • +17b active keeps serving costs reasonable
    • +the widest tooling and fine-tuning ecosystem of any open model
    cons
    • 700m monthly-user gate requiring a separate meta licence
    • derivatives must be named starting with 'llama'
    • rights terminate if you bring ip litigation against meta
    • lmarena submission used a checkpoint meta didn't release
  14. 14

    MiniMax M2.7

    announced as open source, then relicensed after release to forbid commercial use without written permission.

    55/100

    verdicta capable agentic model you cannot ship on. the licence changed after the headlines were written, which is exactly why we read the file rather than the announcement.

    best for
    research and personal projects — and nothing you intend to charge for without an email thread first.
    price
    free for non-commercial only
    pricing note
    any commercial use requires prior written authorisation from minimax; not an osi-compliant licence
    free tier
    yes
    license
    non-commercial
    params
    230b total / 10b active
    context
    ~200k tokens
    commercial use
    needs written permission
    runs on
    multi-gpu node

    the sequence matters. minimax announced 'minimax open sources m2.7' on 12 april, and its previous two models, m2 and m2.5, shipped under standard mit — so the reasonable expectation was more of the same. shortly after release the licence file was revised: commercial use now requires minimax's prior written authorisation, requested by email.

    'commercial use' is defined broadly enough to catch almost anything — any paid product or service, commercial api use, and commercially deployed fine-tunes or derivatives. even authorised users must display 'built with minimax m2.7'. a use-restriction appendix additionally bans military applications, illegal content, disinformation and harassment. the hugging face community flagged it immediately as not osi-compliant, and it isn't.

    which is a shame, because the model is good: 230b total with 10b active, around 200k context, 56.22% on swe-pro matching gpt-5.3-codex, and 57.0% on terminal bench 2. it sits at position 46 on the intelligence index. if the terms were mit it would rank in the top half of this page.

    pros
    • +strong agentic engineering scores for its size
    • +230b total with only 10b active — modest serving cost
    • +free for research and personal use
    cons
    • commercial use needs minimax's prior written permission
    • licence was revised after the 'open sources' announcement
    • predecessors were mit, setting a false expectation
    • mandatory attribution even once authorised
  15. 15

    Kimi K3

    billed as the largest open-weight model ever built, with the weights due tomorrow and a countdown page where the files should be.

    48/100

    verdictlast place is about availability, not merit: as of today there is nothing to download and no terms to read, and this ranking puts the licence first for everything on it.

    best for
    planning rather than building — worth watching if you're choosing a model to standardise on next month.
    price
    unreleased
    pricing note
    weights promised for 27 july 2026; the hugging face repository is a placeholder with no files and no licence, and the model is available only through moonshot's app and api
    free tier
    no
    license
    none published yet
    params
    ~2.8t total (moe)
    context
    1m tokens
    commercial use
    no licence to read
    runs on
    ~1.4tb at 4-bit (estimated)

    moonshot launched k3 on 16 july through its app and api, describing it as the largest open-weight model ever released at roughly 2.8 trillion parameters with a one-million-token context and native multimodal input. the hugging face repository exists as an upcoming-release countdown; on the day we published this, it held no weights and no licence file.

    moonshot's own recent history is the reason not to assume what the terms will be. kimi k2.7 code shipped as 'modified mit', and the modification — an obligation to display the model's name in your product's interface above 100 million monthly users or $20 million in monthly revenue — appears in the licence file and in almost none of the coverage. whatever k3 arrives with should be read in the repository rather than taken from launch day.

    everything else about it is currently unverifiable. there is no public leaderboard entry because there are no weights to evaluate, at least one outlet has flagged an undisclosed hallucination risk pending independent testing, and the circulating hardware figures — around 594gb for the native mxfp4 download and roughly 1.4tb of fast memory at 4-bit — are press estimates with no release to check them against. we will re-rank this entry properly once the files land.

    pros
    • +claimed largest open-weight model ever built at ~2.8t parameters
    • +one-million-token context with native multimodal input
    • +moonshot's previous release shipped genuinely usable terms
    cons
    • no weights published as of this review
    • no licence file exists yet, so the terms are unknown
    • no independent benchmark or leaderboard standing possible
    • hardware requirements are press estimates only

how this ranking was made

every licence claim comes from the licence file in the model's own repository, opened and read. not the launch post, not the model card summary, not a press write-up. where those disagree with the licence, the entry says so and quotes both.

'commercial use' in the spec table is the practical question — can you put this in a product you sell, today, without asking anyone. unrestricted means mit or apache 2.0 with nothing added. anything else is described in the entry.

we rank licence above benchmark. a slightly stronger model you cannot legally ship is worth less than a slightly weaker one you can, and this ranking reflects that deliberately rather than apologetically.

capability standings cite named public leaderboards where a current entry exists — mostly the artificial analysis intelligence index and lmarena — and say 'not verified' where we could not find one. we run no benchmarks ourselves and publish no scores of our own.

self-hosting figures are what the hardware actually needs, and most of these are honest about being out of reach: five entries here need a multi-gpu node or a cluster. where an official vendor figure doesn't exist we say the estimate is a community one.

one licence is genuinely unresolved. nvidia's nemotron 3 ultra links the linux foundation's openmdw-1.1 from its hugging face card, while nvidia's own legal pages describe a custom nemotron open model licence carrying a naming mandate and an indemnification clause that openmdw does not contain. we could not determine which governs. that entry is marked down for the ambiguity, not for either set of terms.

our general methodology and disclosures →
was this useful?