verifier.org

Devstral alternatives

7 tools we tested head to head against Devstral, ranked — and what each one actually does differently.

last reviewed 29 jul 2026 · from our best 8 coding llms ·list curated by Onur Ozcanxin

first — what you'd be leaving

Devstral ranks #5 of 8 in our coding llms testing. the small one is apache 2.0 and the big one stops working for you the month you pass $20m in revenue..

77/100

two models, two licences, and mistral says so plainly — which after codestral from the same company is a meaningful improvement in candour.

why people look for an alternative
  • 123b rights revoked above $20m monthly revenue
  • hard cutoff rather than a graduated licence fee
  • fill-in-the-middle support unconfirmed
  • devstral 2's context length not confirmed

stay with Devstral if devstral small 24b is plain apache 2.0 is the thing you care about most — nothing below beats it on that.

the short version
best alternativeQwen3-Coderrepository-scale agentic work where the context window and the licence both need to be unrestricted.90/100
advertisement
  1. 1

    Qwen3-Coder

    #1 in coding llms · apache 2.0 confirmed in the licence text, a million tokens of context, and now a version small enough to actually run.

    90/100

    verdictthe cleanest licence and the longest context in the category, with a february variant that finally makes it deployable outside a datacentre.

    Qwen3-Coder vs Devstral
     DevstralQwen3-Coder
    pricefree below the thresholdfree (apache 2.0)
    free tieryesyes
    licencemodified mit (123b), apache (24b)apache 2.0
    params123b and 24b480b/35b active; next 80b/3b
    context128k on small262k, 1m via yarn
    commercial usecapped at $20m/month on 123bunrestricted
    runs on~14gb (24b), ~70gb (123b)~240gb, or ~40gb for next

    switch forrepository-scale agentic work where the context window and the licence both need to be unrestricted.

    pros
    • +plain apache 2.0, verified in the licence file
    • +262k native context extending to 1m via yarn
    • +qwen3-coder-next brings it to ~40gb for local use
    • +38.7 on swe-bench pro from the model's own repo
    cons
    • flagship needs ~240gb at 4-bit — server only
    • swe-bench verified figures conflict across sources
    • fill-in-the-middle support unconfirmed for this generation
    • two very different models under one name
  2. 2

    DeepSeek Coder

    #2 in coding llms · the brand is now an api alias — and the models behind it went from a restrictive weights licence to plain mit.

    88/100

    verdictthe strongest swe-bench numbers here under a licence that got more permissive rather than less — with no small variant left for anyone who wanted the original.

    DeepSeek Coder vs Devstral
     DevstralDeepSeek Coder
    pricefree below the thresholdfree (mit)
    free tieryesyes
    licencemodified mit (123b), apache (24b)mit
    params123b and 24b284b/13b active (flash)
    context128k on smallnot confirmed per variant
    commercial usecapped at $20m/month on 123bunrestricted
    runs on~14gb (24b), ~70gb (123b)~140gb at 4-bit

    switch forfrontier-tier coding where the score matters more than running it yourself.

    pros
    • +plain mit on both code and weights, a genuine loosening
    • +80.6% swe-bench verified on v4-pro
    • +frontier-tier coding without a licence negotiation
    • +one-million-token-scale context in the v4 family
    cons
    • no standalone coder models remain — it's an api alias
    • smallest option needs ~140gb at 4-bit
    • nothing runs locally, unlike the original 1.3b-33b line
    • fill-in-the-middle support unconfirmed for v4
  3. 3

    Seed-Coder-8B

    #3 in coding llms · plain mit, native fill-in-the-middle, and it fits in five gigabytes — the practical local pick.

    84/100

    verdictthe best thing here that actually runs on a laptop, with the cleanest licence and the fill-in-the-middle training the frontier models skipped.

    Seed-Coder-8B vs Devstral
     DevstralSeed-Coder-8B
    pricefree below the thresholdfree (mit)
    free tieryesyes
    licencemodified mit (123b), apache (24b)mit
    params123b and 24b8b
    context128k on small32k tokens
    commercial usecapped at $20m/month on 123bunrestricted
    runs on~14gb (24b), ~70gb (123b)~5gb at 4-bit

    switch foreditor completion on your own machine, where latency beats leaderboard position.

    pros
    • +plain mit with no conditions
    • +native fill-in-the-middle and spm training
    • +~5gb at 4-bit — genuinely local
    • +humaneval 84.8 and mbpp 85.2 self-reported
    cons
    • 8b is the only size — no larger variant
    • no update since june 2025
    • 32k context, shortest among the current models here
    • cannot match frontier agentic coders on large repos
    advertisement
  4. 4

    Granite (IBM)

    #4 in coding llms · apache 2.0, native fim, half a million tokens of context — and the dedicated code line is officially deprecated.

    80/100

    verdictthe cleanest licensing story in this ranking and native fim at 5gb — on a product line ibm has folded into its general models.

    Granite (IBM) vs Devstral
     DevstralGranite (IBM)
    pricefree below the thresholdfree (apache 2.0)
    free tieryesyes
    licencemodified mit (123b), apache (24b)apache 2.0
    params123b and 24b3b, 8b, 30b dense
    context128k on small512k in 4.1
    commercial usecapped at $20m/month on 123bunrestricted
    runs on~14gb (24b), ~70gb (123b)~5gb at 4-bit (8b)

    switch forenterprises wanting unambiguous licensing and very long context at sizes that run locally.

    pros
    • +apache 2.0 with licence and documentation in agreement
    • +cryptographically signed weights
    • +native fill-in-the-middle support
    • +512k context in granite 4.1 at ~5gb for the 8b
    cons
    • dedicated granite code line explicitly deprecated
    • coding is now a general-model capability, not a specialism
    • trails dedicated coders on pure code benchmarks
    • legacy code models capped at 4k context
  5. 5

    GLM-4.6

    #6 in coding llms · near-frontier coding under what its badge says is mit — and the licence file itself returns 404.

    72/100

    verdictreportedly close to claude sonnet 4 on real-world coding, with the one thing this ranking insists on — a readable licence — missing.

    GLM-4.6 vs Devstral
     DevstralGLM-4.6
    pricefree below the thresholdfree (mit, unverified)
    free tieryesyes
    licencemodified mit (123b), apache (24b)mit per tag, file unreadable
    params123b and 24b~355b
    context128k on small200k tokens
    commercial usecapped at $20m/month on 123bunrestricted, unverified
    runs on~14gb (24b), ~70gb (123b)~180gb at 4-bit

    switch forhosted agentic coding through claude code or cline-style tooling, where you won't be self-hosting anyway.

    pros
    • +reported near-parity with claude sonnet 4 on cc-bench
    • +200k context, up from 128k in glm-4.5
    • +built for claude code and cline-style integration
    • +mit per the hugging face licence tag
    cons
    • raw licence file returned 404 — text unverified
    • ~180gb at 4-bit, datacentre only
    • superseded by glm-5.1 and 5.2 as general flagships
    • no primary benchmark figure retrievable from zhipu
  6. 6

    StarCoder2

    #7 in coding llms · 600 languages at two gigabytes, under a licence whose conditions you have to pass on to your own users.

    66/100

    verdictstill the best small multi-language coder and two and a half years old — with commercial permission that comes with strings attached to your own terms of service.

    StarCoder2 vs Devstral
     DevstralStarCoder2
    pricefree below the thresholdfree (openrail-m)
    free tieryesyes
    licencemodified mit (123b), apache (24b)bigcode openrail-m v1
    params123b and 24b3b, 7b, 15b
    context128k on small16k, 4k sliding window
    commercial usecapped at $20m/month on 123bpermitted with use restrictions
    runs on~14gb (24b), ~70gb (123b)~2-9gb at 4-bit

    switch forbroad language coverage on small hardware, for teams willing to carry the licence terms downstream.

    pros
    • +600+ programming languages via the stack v2
    • +sizes from ~2gb, the smallest here
    • +native fill-in-the-middle training
    • +15b reported matching codellama-34b
    cons
    • openrail-m restrictions must be passed to your end users
    • disclosure obligation on public-facing generated content
    • no update since february 2024
    • 16k context with 4k sliding-window attention
  7. 7

    Codestral

    #8 in coding llms · downloadable, and forbidden in production — the licence is a flat non-commercial bar, not a threshold.

    58/100

    verdictthe most restrictive licence in this ranking by a distance, on a model whose current iteration may not have open weights at all.

    Codestral vs Devstral
     DevstralCodestral
    pricefree below the thresholdresearch use only
    free tieryesyes
    licencemodified mit (123b), apache (24b)mistral non-production licence
    params123b and 24b22b (open checkpoint)
    context128k on small128k tokens
    commercial usecapped at $20m/month on 123bprohibited without a paid licence
    runs on~14gb (24b), ~70gb (123b)~13gb at 4-bit

    switch forresearch, evaluation and prototypes that will never ship.

    pros
    • +long-established fill-in-the-middle tuned for editors
    • +128k context on the documented model card
    • +22b runs in about 13gb for research use
    • +available as a supported paid api
    cons
    • non-production licence bars commercial use entirely
    • commercial terms require a private negotiation
    • current 25.08 iteration appears to be api-only
    • launch framing understates the prohibition

how these were compared

every tool on this page went through the same test as Devstral — same tasks, same order, scored the same way. the comparison tables are the figures from that testing, not vendor spec sheets.

the coding llms test in full →
was this useful?