verifier.org

Mistral Large 3 alternatives

14 tools we tested head to head against Mistral Large 3, ranked — and what each one actually does differently.

last reviewed 26 jul 2026 · from our best 15 open-weight llms ·list curated by Onur Ozcanxin

first — what you'd be leaving

Mistral Large 3 ranks #3 of 15 in our open-weight llms testing. europe's frontier answer under apache 2.0 — and mistral resisted the temptation to reach for its old research licence..

89/100

the highest-ranked open model on lmarena after gemini, licensed apache 2.0 with nothing bolted on — the main catch is that it's now eight months old.

why people look for an alternative
  • shipped december 2025 — the oldest of the top five here
  • a larger successor is in early access but unreleased
  • 675b total puts self-hosting out of individual reach

stay with Mistral Large 3 if apache 2.0 with no custom mistral terms attached is the thing you care about most — nothing below beats it on that.

the short version
best alternativeGLM-5.2anyone who wants the strongest open model available and needs the legal answer to be boring.95/100
advertisement
  1. 1

    GLM-5.2

    #1 in open-weight llms · the top-ranked open-weight model on the intelligence index, under plain unmodified mit, with a launch post that tells the truth.

    95/100

    verdictfirst on capability and first on licence at the same time, which almost never happens — this is the default open-weight choice right now.

    GLM-5.2 vs Mistral Large 3
     Mistral Large 3GLM-5.2
    pricefree (apache 2.0)free (mit)
    free tieryesyes
    licenseapache 2.0mit
    params675b total / 41b active753b total / ~40b active
    context256k tokens1m tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb cluster8x h200 node

    switch foranyone who wants the strongest open model available and needs the legal answer to be boring.

    pros
    • +plain unmodified mit — no traps of any kind
    • +top-ranked open-weight model on the artificial analysis index
    • +one-million-token context
    • +launch claims and licence file actually agree
    cons
    • roughly 750gb of gpu memory at fp8 — a full node
    • active-parameter count only confirmed via secondary sources
    • no published vendor statement of its weaknesses
  2. 2

    DeepSeek V4 Pro

    #2 in open-weight llms · 1.6 trillion parameters of mit-licensed coding ability, still shipping as a preview three months after its stable release was due.

    92/100

    verdictarguably the best coding model you can legally do anything with — carrying a 'preview' label that deepseek has not removed on the schedule it set itself.

    DeepSeek V4 Pro vs Mistral Large 3
     Mistral Large 3DeepSeek V4 Pro
    pricefree (apache 2.0)free (mit)
    free tieryesyes
    licenseapache 2.0mit
    params675b total / 41b active1.6t total / 49b active
    context256k tokens1m tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb clustermulti-node cluster

    switch forcoding and reasoning workloads where you want frontier results without a frontier bill or a licence review.

    pros
    • +plain mit at genuine frontier coding capability
    • +80.6% on swe-bench verified per vendor reporting
    • +one-million-token context
    • +cheapest frontier hosting of any model here
    cons
    • still labelled preview; promised stable release never shipped
    • roughly 1.9tb at fp16 — multi-node self-hosting
    • tool-use and multimodal breadth trail closed rivals
  3. 3

    Kimi K2.7 Code

    #4 in open-weight llms · a trillion-parameter agentic coder under 'modified mit' — where the modification only bites if you get very big.

    86/100

    verdictexcellent at the job it was built for, with a licence quirk almost nobody will hit — but which press coverage summarising it as 'mit' never mentions.

    Kimi K2.7 Code vs Mistral Large 3
     Mistral Large 3Kimi K2.7 Code
    pricefree (apache 2.0)free (modified mit)
    free tieryesyes
    licenseapache 2.0modified mit
    params675b total / 41b active1t total / 32b active
    context256k tokens256k tokens
    commercial useunrestrictedattribution above scale
    runs on8x80gb cluster2x a100 at int4

    switch forlong-running agentic engineering work, for anyone comfortably below the attribution thresholds.

    pros
    • +attribution clause only triggers at 100m users or $20m monthly revenue
    • +purpose-built for long-horizon agentic engineering
    • +~30% fewer thinking tokens than its predecessor
    • +int4 deployment reportedly fits two a100 80gb cards
    cons
    • 'modified mit' is described as plain mit almost everywhere
    • trails frontier closed models on its vendor's own benchmarks
    • position 31 on the artificial analysis index
    • 256k context, against 1m for the top three
    advertisement
  4. 4

    Inkling

    #5 in open-weight llms · thinking machines' first model — apache 2.0, multimodal, and unusually honest about not being the best.

    84/100

    verdictthe strongest open-weight model from a us lab, built for customisation rather than benchmarks — and its own model card says plainly that it isn't the strongest overall.

    Inkling vs Mistral Large 3
     Mistral Large 3Inkling
    pricefree (apache 2.0)free (apache 2.0)
    free tieryesyes
    licenseapache 2.0apache 2.0
    params675b total / 41b active975b total / 41b active
    context256k tokens1m tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb cluster8x80gb+ cluster

    switch forteams who want a permissive multimodal base to fine-tune rather than a leaderboard winner to deploy.

    pros
    • +apache 2.0 from a major new us lab
    • +genuinely multimodal — text, image, audio and video training
    • +up to 1m context from downloaded weights
    • +unusually candid about its own limitations
    cons
    • licence confirmed from the model card, not the raw repo file
    • behind deepseek v4 pro and kimi k2.6 on the intelligence index
    • 975b total — enterprise-scale hosting only
    • tinker api caps context at 256k, lower than the weights allow
  5. 5

    gpt-oss-120b

    #6 in open-weight llms · the only genuinely capable model here that fits on a single 80gb card, under clean apache 2.0.

    82/100

    verdictnearly a year old and still the most practical self-host on this list — one card, no licence questions, real agentic ability.

    gpt-oss-120b vs Mistral Large 3
     Mistral Large 3gpt-oss-120b
    pricefree (apache 2.0)free (apache 2.0)
    free tieryesyes
    licenseapache 2.0apache 2.0
    params675b total / 41b active117b total / 5.1b active
    context256k tokens128k tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb clustersingle 80gb gpu

    switch forteams who want to actually own the deployment rather than rent it, on hardware they can afford.

    pros
    • +unmodified apache 2.0 with no layered openai terms
    • +runs on a single 80gb gpu at mxfp4
    • +5.1b active path makes cpu inference viable
    • +strong agentic tool-calling with configurable reasoning effort
    cons
    • text-only in a multimodal field
    • 128k context, well below the 2026 frontier
    • no successor since august 2025
    • passed on the intelligence index by newer open releases
  6. 6

    DeepSeek V4 Flash

    #7 in open-weight llms · a separate model rather than a trim of v4 pro — a fifth the size, same mit licence, a twelfth the hosted price.

    81/100

    verdictthe best price-to-capability ratio in the category, and widely mistaken for a variant of v4 pro when it's its own release with its own weights.

    DeepSeek V4 Flash vs Mistral Large 3
     Mistral Large 3DeepSeek V4 Flash
    pricefree (apache 2.0)free (mit)
    free tieryesyes
    licenseapache 2.0mit
    params675b total / 41b active284b total / 13b active
    context256k tokens1m tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb clustersingle 8x80gb node

    switch forhigh-volume work that doesn't need pro-tier reasoning but does need a licence you can ignore.

    pros
    • +plain mit, same as v4 pro
    • +cheapest frontier-adjacent hosted rate in the market
    • +13b active — far more self-hostable than the pro tier
    • +a fifth the total size for a fraction of the capability gap
    cons
    • routinely misdescribed as a variant of v4 pro
    • 40.3% on the intelligence index — clearly below the top tier
    • inherits v4's preview-status uncertainty
  7. 7

    Gemma 4

    #8 in open-weight llms · the year's biggest licensing improvement — google dropped its custom terms and shipped this one under real apache 2.0.

    79/100

    verdictthree generations of custom gemma terms replaced by plain apache 2.0, on a family built to run on hardware you already own — the caveats are housekeeping rather than substance.

    Gemma 4 vs Mistral Large 3
     Mistral Large 3Gemma 4
    pricefree (apache 2.0)free (apache 2.0)
    free tieryesyes
    licenseapache 2.0apache 2.0
    params675b total / 41b active31b dense / 26b moe
    context256k tokens256k tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb clustersingle consumer gpu

    switch foron-device and laptop-class multimodal work where the model has to run locally and legally.

    pros
    • +genuine apache 2.0, replacing three generations of custom terms
    • +multimodal variants that run on laptops and edge devices
    • +256k context and 140+ languages
    • +31b dense fits a single consumer gpu when quantised
    cons
    • repos still gated behind an accept-terms click-through
    • google's old custom terms page is still live and unclarified
    • training data cutoff of january 2025
  8. 8

    Qwen3.6-35B-A3B

    #9 in open-weight llms · genuinely apache 2.0 and genuinely small — and constantly confused with the closed flagship sharing its brand.

    77/100

    verdictan unusually strong small model under clean apache 2.0 — the trap is definitional, because the 'qwen' name also covers a closed api-only flagship.

    Qwen3.6-35B-A3B vs Mistral Large 3
     Mistral Large 3Qwen3.6-35B-A3B
    pricefree (apache 2.0)free (apache 2.0)
    free tieryesyes
    licenseapache 2.0apache 2.0
    params675b total / 41b active35b total / 3b active
    context256k tokensnot confirmed
    commercial useunrestrictedunrestricted
    runs on8x80gb clustersingle high-end gpu

    switch foragentic coding on modest hardware, for teams who want alibaba's engineering without alibaba's api.

    pros
    • +standard unmodified apache 2.0
    • +3b active — runs on one prosumer gpu
    • +vendor-reported 73.4 on swe-bench for its size
    • +combined thinking and non-thinking modes
    cons
    • constantly conflated with the closed api-only qwen3.7-max
    • 35b total means limited knowledge breadth
    • context length not confirmed from a primary source
    • benchmark figures are vendor-reported only
  9. 9

    NVIDIA Nemotron 3 Ultra

    #10 in open-weight llms · a serious frontier model whose licence nvidia's own two sets of pages cannot agree on.

    74/100

    verdictcapable, long-context and well-engineered — marked down purely because we could not establish, from nvidia's own sources, what you are agreeing to.

    NVIDIA Nemotron 3 Ultra vs Mistral Large 3
     Mistral Large 3NVIDIA Nemotron 3 Ultra
    pricefree (apache 2.0)free (licence disputed)
    free tieryesyes
    licenseapache 2.0openmdw-1.1 or custom — disputed
    params675b total / 41b active550b total / 55b active
    context256k tokens1m tokens
    commercial useunrestrictedunresolved
    runs on8x80gb clustermulti-node cluster

    switch forenterprises with legal teams who can get nvidia to say which licence actually governs before deployment.

    pros
    • +550b/55b active with a one-million-token context
    • +built for long-horizon agentic reasoning and tool-calling
    • +openmdw-1.1, if it governs, is genuinely permissive
    cons
    • nvidia's own pages name two different licences for it
    • the custom licence adds a naming mandate and indemnification
    • 550b total — enterprise-only self-hosting
    • no verified public leaderboard standing found
  10. 10

    OLMo 3

    #11 in open-weight llms · the only model here where 'fully open' survives contact with the evidence — weights, training data, code and checkpoints.

    72/100

    verdictevery other entry calls itself open and means the weights. ai2 publishes the training data too — which is worth more than the benchmark places it gives up.

    OLMo 3 vs Mistral Large 3
     Mistral Large 3OLMo 3
    pricefree (apache 2.0)free (apache 2.0)
    free tieryesyes
    licenseapache 2.0apache 2.0
    params675b total / 41b active7b and 32b dense
    context256k tokens32k tokens
    commercial useunrestrictedunrestricted
    runs on8x80gb clustersingle consumer gpu

    switch forresearch, auditing, and anyone who needs to know what a model was actually trained on.

    pros
    • +training data, code and checkpoints published, not just weights
    • +clean apache 2.0
    • +7b variant runs on a consumer gpu
    • +think and rlzero variants built for reasoning research
    cons
    • 32k context — the shortest in this ranking
    • benchmark ceiling well below the frontier open models
    • no current leaderboard standing found
    • released late 2025 with no 2026 update
+ 4 more tested, not detailed here
we ranked 15 open-weight llms in total. the 4 that didn't make this page are written up in the full ranking →

how these were compared

every tool on this page went through the same test as Mistral Large 3 — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.

the open-weight llms test in full →
was this useful?