verifier.org

best 20 ai image models

ranked on prompt adherence, text rendering, editing, licence terms you can actually ship on, and real cost per image.

last reviewed 22 jul 2026 · 20 tools tested ·list curated by Onur Ozcanxin

the short version
best overallGPT Image 2typography, diagrams, and any prompt with a dozen constraints that all have to survive95/100runner-upReve 2.1dense compositions, posters, and anything that needs to stay coherent at very large sizes90/100best free optionHiDream-O1-Imageself-hosted commercial products that need weights you can actually ship on80/100

raw fidelity stopped being the differentiator some time ago. everything in the top half of this list makes a technically clean image. what separates them now is whether the model does what the prompt actually said, whether it can render legible text, whether you can edit one thing without the rest of the frame drifting, and what a thousand images actually cost.

licensing is the axis people keep underweighting, and 2026 made it worse. several of the most-downloaded 'open' models — flux.2 dev, flux.2 klein 9b, ideogram 4 — ship weights under non-commercial licences. teams benchmark them, like them, and ship into breach. the two genuinely permissive frontier-adjacent open models right now are nvidia's cosmos3 and hidream-o1, which is a completely different cast than a year ago.

one structural note on price: routing google's models through an aggregator like fal costs roughly 12–25% more than going direct to the gemini api, and google also offers 50% batch pricing that aggregators don't expose. the convenience of one unified api is real, but so is the markup.

two models are deliberately absent. meta's muse image would rank around third on public leaderboards but is consumer-app-only with no api. alibaba's qwen-image 3.0 shipped on 21 july with no weights, no licence, no paper and no benchmarks — nothing independently testable. we'll rank both when you can actually buy them.

advertisement
  1. 1

    GPT Image 2

    the only model that leads both generation and editing

    95/100

    verdictthe best model in the category on every measure that can be measured — and priced so that you should still draft on something else.

    best for
    typography, diagrams, and any prompt with a dozen constraints that all have to survive
    price
    $0.006 low / $0.053 medium / $0.211 high, at 1024×1024
    pricing note
    billed per image by quality tier; a 35× spread between low and high makes forecasting genuinely hard
    free tier
    no
    access
    closed api + chatgpt
    max resolution
    3840px max edge, 3:1 aspect limit
    text rendering
    best in class
    editing
    inpaint, outpaint, multi-reference
    license
    commercial via api terms

    gpt image 2 is the only model that sits at number one on artificial analysis for both text-to-image and editing, and number one on lmarena with the largest vote pool of any entrant. that combination has not happened before; usually the aesthetic leader and the instruction-following leader are different products.

    its signature strength is fine typography. multi-line legible text, correct spelling, sensible kerning, text that sits properly inside a shape — for infographics, packaging mockups and social assets with real copy, it is routinely the only model that gets it right on the first generation.

    the reason it isn't an automatic default is the price curve. high quality at 1024×1024 is $0.211 per image, roughly three times nano banana 2 and six times its lite sibling. the low tier is $0.006, which is competitive, but the gap between them is 35× — so a pipeline that quietly defaults to high will produce a bill nobody forecast.

    one editing caveat straight from openai's own docs: masking is prompt-based rather than pixel-exact, so it may not follow the precise shape you painted. it is still the best inpainting available; it just isn't the surgical tool the word 'mask' implies.

    pros
    • +#1 on artificial analysis for generation and editing, and #1 on lmarena
    • +best fine typography and multi-line text of any model
    • +holds long prompts with many simultaneous constraints
    • +quality tiers let you draft cheap and pay only for keepers
    cons
    • $0.211 per high-quality image is the priciest mainstream option
    • 35× price spread between tiers makes cost forecasting hard
    • masking is prompt-based, not pixel-exact
  2. 2

    Reve 2.1

    the independent lab that reasons about layout before it draws

    90/100

    verdictsecond on both public leaderboards and the only model here with genuine 16-megapixel native output — the strongest thing built outside a hyperscaler.

    best for
    dense compositions, posters, and anything that needs to stay coherent at very large sizes
    price
    ~$0.04 / image on fal's reve endpoint
    pricing note
    reve has not published 2.1-specific pricing — the $0.04 figure is fal's unversioned reve endpoint, so confirm before budgeting
    free tier
    yes
    access
    closed api + web app
    max resolution
    4096×4096 native (16mp)
    text rendering
    strong, incl. non-latin
    editing
    instruction edit + 8-image remix
    license
    commercial via api terms

    reve's architecture reasons about structure and hierarchy before rendering, rather than denoising its way to a composition and hoping the layout survives. in practice that shows up exactly where you would expect: dense posters, editorial layouts, and prompts with many elements that all need their own space stay coherent instead of collapsing into mush.

    it renders at 4096×4096 natively — sixteen megapixels, not an upscale. nothing else on this list does that, and for print work or anything that gets cropped hard afterwards it removes a whole post-processing step.

    it sits second on both artificial analysis and lmarena, which is a genuinely surprising result for a company this size. the caveat worth knowing is that its lmarena vote count is a fraction of gpt image 2's, so the ranking is real but less settled — expect it to move more than the top spot does.

    editing is covered by two modes: instruction-based edit on a single image, and remix across up to eight references. the missing piece is pricing transparency — reve does not publish a per-image rate for 2.1 anywhere we could find, which is an odd gap for a product this good.

    pros
    • +#2 on both artificial analysis and lmarena
    • +4096×4096 native output, the highest here by a wide margin
    • +layout reasoning keeps dense compositions coherent
    • +remix accepts up to eight reference images
    cons
    • no published pricing for 2.1 specifically
    • leaderboard position rests on far fewer votes than the leader's
    • smaller ecosystem and thinner documentation than the hyperscalers
  3. 3

    Nano Banana 2

    google's gemini 3.1 flash image — the best default for production work

    89/100

    verdictnot the top of any leaderboard, and still the model we'd default to — the balance of price, speed, resolution and character consistency is better than anything above it.

    best for
    high-volume commercial pipelines that need consistent characters and predictable cost
    price
    $0.067 / image at 1K direct from google
    pricing note
    $0.045 at 0.5K, $0.101 at 2K, $0.151 at 4K; fal charges ~19% more; google offers 50% batch pricing that aggregators don't expose
    free tier
    yes
    access
    closed api (gemini, vertex) + apps
    max resolution
    512px to 4K
    text rendering
    strong, incl. in-image translation
    editing
    excellent targeted editing
    license
    commercial; synthid watermarked

    the naming is a mess, so to be clear: nano banana 2 is gemini 3.1 flash image, released in february. it is not a successor to nano banana pro in the quality sense — it is the fast line catching up, which it comprehensively has.

    its standout capability is subject consistency: up to five characters and fourteen objects held stable across generations. for anyone producing a series — product shots, a character across a storyboard, a campaign that has to look like one campaign — that is worth more than a few elo points, and it is class-leading.

    it ranks third on artificial analysis' editing board, ahead of the more expensive nano banana pro, and does 4K, web grounding and in-image translation. at $0.067 per image at 1K it costs less than a third of gpt image 2 at high quality.

    where it loses is raw aesthetic preference — head to head, voters pick gpt image 2 by a wide margin. if a single hero image has to be the best possible, this isn't it. if you need ten thousand good ones, it is.

    pros
    • +best price/quality/speed balance on the market
    • +five-character, fourteen-object consistency is class-leading
    • +#3 on the editing leaderboard, ahead of nano banana pro
    • +4K output, web grounding, and 50% batch pricing direct from google
    cons
    • clearly behind gpt image 2 on aesthetic preference
    • the nano banana naming makes version choice genuinely confusing
    • ~19% markup if you route through fal instead of google
    advertisement
  4. 4

    Nano Banana 2 Lite

    the budget tier that outscores the flagship it's cheaper than

    87/100

    verdictscores above full nano banana 2 on artificial analysis at half the price and nearly three times the speed — the value story of 2026.

    best for
    high-volume generation where 1024px is enough and latency matters
    price
    $0.034 / image at 1K direct from google
    pricing note
    on fal it is billed by token at $37.50/1M image-out with fixed 1024×1024 output — roughly 25% more than going direct
    free tier
    yes
    access
    closed api (gemini, vertex)
    max resolution
    1024×1024 via fal
    text rendering
    strong for the price
    editing
    supported, flash-tier quality
    license
    commercial; synthid watermarked

    this is the entry that breaks the usual assumption. released at the end of june, nano banana 2 lite posts a higher artificial analysis elo than the full nano banana 2 it sits beneath in google's own lineup, at $0.034 per image against $0.067, and generates in about four seconds — roughly 2.7× faster.

    cheap tiers being strictly worse used to be a safe rule. it isn't any more, and if you are generating at volume you should benchmark this before paying for anything above it.

    the constraint is resolution. through fal it is pinned to 1024×1024, so anything needing 2K or 4K has to move up the family. it also inherits the flash line's aesthetic — competent and slightly literal rather than striking.

    the billing model on fal is token-based rather than per-image, which makes cost modelling more annoying than it should be. going direct to google is both cheaper and easier to reason about.

    pros
    • +outscores full nano banana 2 on artificial analysis
    • +half the price of nano banana 2, ~2.7× faster
    • +$0.034/image is frontier-adjacent quality at budget cost
    • +same google reliability, sdks and batch discount
    cons
    • fixed 1024×1024 output through fal — no 2K or 4K
    • token-based billing on aggregators is awkward to forecast
    • aesthetic is literal rather than striking
  5. 5

    MAI-Image-2.5

    microsoft's own model, strong scores, awkward to buy

    84/100

    verdictthird on artificial analysis and genuinely excellent at text and stylised art — held back almost entirely by how hard it is to buy if you aren't already on azure.

    best for
    teams already inside azure or microsoft foundry
    price
    token-billed: $47 / 1M image-output tokens
    pricing note
    microsoft publishes token rates only, no per-image figure; the Flash variant is $33/1M image-out
    free tier
    no
    access
    closed api (foundry, openrouter)
    max resolution
    not published
    text rendering
    excellent
    editing
    control-with-preservation editing
    license
    commercial via api terms

    released in june, mai-image-2.5 is microsoft's own image model rather than a rebadged openai one, and it is good: third on artificial analysis text-to-image, fourth on editing, sixth on lmarena. the jump over the previous generation was largest in text rendering and in cartoon, anime and fantasy styles.

    its editing approach — microsoft calls it control with preservation — is aimed squarely at the failure mode where a targeted change quietly redraws the rest of the frame, and it handles that well.

    distribution is the problem. it is available through microsoft foundry, the mai playground and openrouter, and it is shipping inside powerpoint and onedrive, but there is no clean consumer-facing api story and microsoft publishes token rates rather than a per-image price. working out what a thousand images costs takes a spreadsheet.

    if you are already an azure shop, this is a strong and under-discussed option. if you aren't, the friction is real and the models above it are easier to adopt.

    pros
    • +#3 on artificial analysis text-to-image, #4 on editing
    • +large generational jump in text rendering and stylised art
    • +editing preserves the rest of the frame well
    • +already embedded in powerpoint and onedrive
    cons
    • token-only pricing with no published per-image cost
    • distribution is effectively azure-first
    • little community tooling or ecosystem
  6. 6

    Nano Banana Pro

    google's gemini 3 pro image — the accuracy specialist, now overpriced

    83/100

    verdictstill the best model for grounded, factually accurate imagery — but its own cheaper stablemate now beats it at editing, which makes the price hard to defend.

    best for
    infographics and anything where factual accuracy in the image matters more than cost
    price
    $0.134 / image at 1K–2K, $0.24 at 4K
    pricing note
    fal charges $0.15 and $0.30, plus $0.015 when web search is used
    free tier
    yes
    access
    closed api (gemini, vertex) + apps
    max resolution
    4K
    text rendering
    strong
    editing
    strong, #5 on editing board
    license
    commercial; synthid watermarked

    nano banana pro is gemini 3 pro image, and its distinguishing trick is that it can search the web mid-generation. for infographics, diagrams with real data, or anything where the picture makes a factual claim, that grounding produces results the others simply can't.

    it was google's quality flagship from november until february, and google's own guidance still points here for high-fidelity work requiring maximum factual accuracy.

    the problem is internal competition. nano banana 2 ranks above it on the editing leaderboard at half the price, and nano banana 2 lite outscores it on price-performance by a wider margin still. unless you specifically need the web grounding or 4K at this tier, the cheaper siblings are the better buy.

    it is also the slow one in the family. for interactive or high-volume work the latency is noticeable.

    pros
    • +web grounding produces genuinely factual imagery
    • +best-in-family world knowledge for infographics
    • +native 4K output
    • +google reliability, sdks and batch pricing
    cons
    • beaten on editing by nano banana 2 at half the price
    • slowest of the nano banana family
    • $0.24 at 4K is expensive for what it now delivers
  7. 7

    Seedream 5.0 Pro

    the best non-latin text rendering available anywhere

    82/100

    verdictif your text isn't in english, this is the model — ten-plus languages rendered natively, and nothing else comes close on cjk.

    best for
    multilingual design work, especially cjk and other non-latin scripts
    price
    $0.0675 / image up to 1536×1536
    pricing note
    $0.135 from there to 2048×2048; available via volcano engine, byteplus and fal
    free tier
    no
    access
    closed api (volcano, byteplus, fal)
    max resolution
    2048×2048, ~3K long edge
    text rendering
    best for non-latin scripts
    editing
    layered + precise interactive edit
    license
    commercial via api terms

    bytedance released seedream 5.0 pro two weeks ago, and its differentiator is unambiguous: native multilingual text rendering across more than ten languages. western models fall apart on chinese, japanese and korean characters in a way that is obvious to anyone who reads them. this one doesn't.

    it is also built for information density — charts, dense infographics, layouts with a lot of small type — and supports multi-layer separation and precise interactive editing, which makes it unusually practical for design handoff rather than just generation.

    at $0.0675 per image up to 1536², it is priced competitively against nano banana 2 while doing something none of the models above it can.

    two caveats. it is brand new and does not yet appear in either leaderboard's top ten, so our placement leans on capability rather than public voting. and for western enterprises there is procurement friction around a bytedance model that is worth raising internally before it becomes a surprise.

    pros
    • +best non-latin and cjk text rendering by a clear margin
    • +ten-plus languages rendered natively
    • +strong on dense information graphics and layered output
    • +competitively priced at $0.0675
    cons
    • absent from both leaderboard top tens despite the capability
    • bytedance procurement friction for some western buyers
    • 2048×2048 ceiling, lower than its own lite sibling
  8. 8

    HiDream-O1-Image

    8b parameters, MIT licence, genuinely shippable

    80/100

    verdictthe best open model you can legally use in a commercial product — and it beats models seven times its size.

    best for
    self-hosted commercial products that need weights you can actually ship on
    price
    free — open weights
    pricing note
    MIT licence, no revenue cap, no commercial restriction; you pay only for gpus
    free tier
    yes
    access
    open weights, self-hosted
    max resolution
    model-dependent, typically 1–2K
    text rendering
    good
    editing
    via community tooling
    license
    MIT — fully commercial

    hidream-o1 is pixel-native: no vae, no separate text encoder, and it reasons before it draws. at 8 billion parameters it trades punches with flux.2 dev at 32 billion, and its 1.5 revision sits fourth on artificial analysis.

    the reason it ranks this high is the licence. it is MIT. no revenue cap, no non-commercial clause, no 'contact us above 100 million users'. in a year where most of the open-weight headline releases turned out to be non-commercial, that is the differentiator that actually matters.

    self-hosting still means gpu capacity, ops, and no vendor to call at 2am. this is the right pick when you need control, data residency, or per-image economics that only work if inference is yours — not when you want the shortest path to an image.

    ecosystem is the weak point. the lora and controlnet tooling that grew up around flux does not exist here in the same depth, so expect more integration work.

    pros
    • +MIT licence — unrestricted commercial use
    • +competitive with models 4× its parameter count
    • +#4 on artificial analysis for the 1.5 revision
    • +pixel-native architecture runs without a separate text encoder
    cons
    • self-hosting means real gpu and ops cost
    • much thinner tooling ecosystem than flux
    • no hosted first-party api to fall back on
  9. 9

    FLUX.2 [pro]

    ten reference images and the deepest tooling ecosystem in the category

    79/100

    verdictno longer leads any leaderboard, but nothing else takes ten reference images — for product and brand work that single capability still wins.

    best for
    brand and product consistency across large volumes of images
    price
    $0.03 first megapixel + $0.015 per extra megapixel
    pricing note
    fal matches black forest labs' direct pricing here — no aggregator markup, unlike the google models
    free tier
    no
    access
    closed api (bfl, fal)
    max resolution
    4mp editing
    text rendering
    good
    editing
    10-reference composition, 4mp
    license
    commercial via api; weights are not

    flux.2 pro accepts up to ten reference images and lets you address them individually in the prompt. for keeping a product, a face or a brand style consistent across hundreds of generations, that is a different class of control from a single style reference, and it is the main reason this model is still in the top ten.

    it edits at four megapixels, handles photorealism well, and its pricing is honest — $0.03 for the first megapixel and $0.015 for each additional one, matching black forest labs' direct rate rather than marking it up.

    it has, however, fallen out of the leaderboard top ten entirely. a year ago flux was the answer; today it is a specialist with an unusually deep toolchain around it.

    the family is also a licensing minefield, and this is the entry to read carefully. the api models are fine — you are buying hosted inference. the downloadable ones are not: flux.2 dev and flux.2 klein 9b both ship under the flux non-commercial licence despite launch messaging that reads otherwise. generated outputs can be used commercially; the weights cannot.

    pros
    • +up to ten reference images, individually addressable
    • +best-in-class product and brand consistency
    • +4mp editing and strong photorealism
    • +no aggregator markup — fal matches bfl direct pricing
    cons
    • no longer in the leaderboard top ten
    • the wider flux family's licensing traps catch a lot of teams
    • megapixel billing makes large images add up fast
  10. 10

    NVIDIA Cosmos3

    the most completely open release in the category

    77/100

    verdictthe top open-weight model on artificial analysis, released with datasets and training recipes — if provenance matters to you, nothing else is close.

    best for
    research, fine-tuning, and teams who need to see what the model was trained on
    price
    free — open weights
    pricing note
    OpenMDW-1.1 licence with commercial rights; checkpoints, six datasets and training recipes all published
    free tier
    yes
    access
    open weights, self-hosted
    max resolution
    model-dependent
    text rendering
    adequate
    editing
    via community tooling
    license
    OpenMDW-1.1 — commercial ok

    cosmos3 is the highest-ranked open-weight model on artificial analysis, and nvidia released it properly: checkpoints, six datasets, and the training recipes. most 'open' releases in 2026 mean a weights file and a licence that forbids using it. this one means what the word used to mean.

    the openmdw-1.1 licence grants commercial rights, which together with hidream's MIT makes these two the only genuinely shippable open models near the frontier.

    it is also agentic — it plans and revises rather than denoising straight from the prompt — which is where the whole category moved this year.

    practically, it is a heavier lift than hidream: bigger, more infrastructure, and aimed at teams who want to fine-tune rather than teams who want to call an api. if you need auditable training data for a regulated client, the dataset release is worth the effort on its own.

    pros
    • +highest-ranked open-weight model on artificial analysis
    • +openmdw-1.1 licence permits commercial use
    • +datasets and training recipes published, not just weights
    • +agentic planning rather than straight denoising
    cons
    • heavier infrastructure requirement than hidream-o1
    • aimed at fine-tuners more than api consumers
    • no first-party hosted endpoint
  11. 11

    Ideogram 4.0

    still the typography specialist, with a licence you must read

    76/100

    verdictstill the best model for stylised lettering, and the cheapest good option per megapixel — but the downloadable weights are non-commercial, whatever the press coverage says.

    best for
    posters, packaging, signage — anything where the words are the design
    price
    $0.00525 / MP turbo, $0.0105 balanced, $0.0175 quality
    pricing note
    a 2048×2048 image runs $0.021 to $0.07 depending on tier
    free tier
    yes
    access
    closed api + web + weights
    max resolution
    2K native
    text rendering
    best for stylised type
    editing
    inpaint, remix, layout boxes
    license
    weights non-commercial; api commercial

    ideogram built its identity on rendering text correctly and it still leads on typography specifically: dense small type, multi-word headlines, lettering that follows a curve or sits properly inside a shape. it is also number one open-weight model on designarena.

    version 4 added bounding-box layout control and a structured json prompting interface, which is genuinely differentiated — if you are generating hundreds of assets that each carry a headline in a fixed position, that turns a creative tool into a pipeline component.

    at $0.00525 per megapixel on the turbo tier it is also among the cheapest credible models here.

    now the important part. ideogram 4 is widely reported as open weights under a commercial licence. the model cards on hugging face say otherwise — both the fp8 and nf4 releases carry an ideogram 4 non-commercial agreement, and fal's own derivative states that it inherits that non-commercial term. treat the weights as non-commercial and the hosted api as the commercial path, and get written confirmation from ideogram before you plan otherwise.

    pros
    • +category leader for stylised typography and lettering
    • +bounding-box layout control and json prompting
    • +cheapest quality tier per megapixel on this list
    • +generous free tier for evaluation
    cons
    • downloadable weights are non-commercial despite press claims
    • general-purpose quality trails the top five
    • 2K native ceiling
  12. 12

    Recraft V4.1

    the only one that outputs real vectors

    75/100

    verdictgenuine editable svg output with no real competitor — for design systems it isn't the best model, it's the only one.

    best for
    icon sets, logos, and design-system work that has to stay editable
    price
    ~$0.035 / image (V4.1); V4 Pro ~$0.25
    pricing note
    recraft's own pricing page doesn't render figures — these come from aggregators, so confirm before budgeting
    free tier
    yes
    access
    closed api + web app
    max resolution
    vector — resolution independent
    text rendering
    good in vector layouts
    editing
    vector editing + style sets
    license
    commercial on paid plans

    recraft generates actual vector graphics: editable svg, not a raster image auto-traced afterwards. for icon sets, logo exploration and spot illustration that go straight into figma and stay editable, that is a category difference rather than a feature.

    brand style sets are the other strong idea — train a style from references and every subsequent generation stays on-brand, which is the consistency problem that makes most image models unusable for production design systems.

    v4.1 utility pro sits tenth on artificial analysis, which is a respectable showing for a tool this specialised.

    photoreal work is not its strength and it doesn't pretend otherwise. it is a design tool that happens to use generative models and should be judged that way. pricing is also frustratingly hard to pin down — recraft's own pricing page didn't render figures for us, so the numbers here come from aggregators.

    pros
    • +true editable svg output — effectively unique
    • +brand style sets give real cross-asset consistency
    • +#10 on artificial analysis despite the narrow focus
    • +strong at icons, infographics and spot illustration
    cons
    • photoreal generation well behind the leaders
    • first-party pricing is hard to verify
    • narrower use case than a general model
  13. 13

    Krea 2 Turbo

    the best aesthetics-per-dollar, built around style control

    74/100

    verdictthe cheapest route to a genuinely good-looking image, with the most thoughtful style controls in the category.

    best for
    art direction, moodboarding, and iterating on a look rather than a fact
    price
    $0.008 / megapixel
    pricing note
    one of the cheapest quality models available; krea 2 turbo generates in roughly two seconds
    free tier
    yes
    access
    closed api + web app
    max resolution
    megapixel-billed, flexible
    text rendering
    weak
    editing
    style refs, loras, sliders
    license
    commercial on paid plans

    krea's whole thesis is style control. moodboards, multiple style references with adjustable influence weights, lora support, and generative sliders for intensity, complexity and movement. it is built for people who know what they want it to look like and are tired of describing it in adjectives.

    krea 2 turbo generates in about two seconds at $0.008 per megapixel, which makes it viable for the kind of rapid iteration that is prohibitively expensive on gpt image 2.

    it is not a text-rendering or factual-accuracy model, and it doesn't try to be. put a headline in the prompt and you will be disappointed.

    one thing to be careful with: krea markets itself as sixth globally on artificial analysis and first among independent labs. we could not corroborate the sixth placement on the leaderboard we pulled, where it sits outside the top ten. the model is good and the price is excellent — the self-reported ranking is the part we'd discount.

    pros
    • +$0.008/mp is among the cheapest quality options anywhere
    • +best style-reference and moodboard controls in the category
    • +generative sliders for intensity, complexity and movement
    • +roughly two-second generation
    cons
    • weak text rendering — not a typography tool
    • self-reported leaderboard ranking we couldn't verify
    • not aimed at factual or informational imagery
  14. 14

    Seedream 5.0 Lite

    higher resolution than the pro model, at half the price

    72/100

    verdictan odd product — cheaper and higher-resolution than the pro tier above it, and worth benchmarking before you assume pro is the upgrade.

    best for
    cheap high-resolution generation, especially with multilingual text
    price
    $0.035 / image
    pricing note
    outputs up to 3072×3072 — counterintuitively higher than seedream 5.0 pro's 2048 ceiling
    free tier
    no
    access
    closed api (volcano, byteplus, fal)
    max resolution
    3072×3072
    text rendering
    strong, multilingual
    editing
    basic relative to pro
    license
    commercial via api terms

    seedream 5.0 lite outputs at up to 3072×3072, which is larger than seedream 5.0 pro's 2048 ceiling, at roughly half the price. that is not how product tiers normally work, and it is worth knowing before you default to the pro model.

    it carries the same architecture combining deep thinking, web search and generation, and inherits the family's strong multilingual text rendering.

    what you give up is the pro model's layer separation, its precise interactive editing, and some quality at the top end. for a lot of production work that trade is fine.

    as with the pro model, it is new enough that public leaderboards haven't placed it, and bytedance procurement questions apply equally.

    pros
    • +3072×3072 output — higher than the pro tier
    • +$0.035 per image is genuinely cheap for the resolution
    • +inherits strong multilingual text rendering
    • +unified thinking, search and generation architecture
    cons
    • no layer separation or precise interactive editing
    • unplaced on public leaderboards
    • same bytedance procurement friction as the pro model
  15. 15

    Grok Imagine

    good editing, permissive content policy, and a billing practice you should know about

    69/100

    verdicta solid editing model with a notably permissive content policy — and the only one here that bills you for generations it refuses to produce.

    best for
    editing work and subjects other models refuse
    price
    $0.02 standard, $0.05 at 1K / $0.07 at 2K on the quality tier
    pricing note
    fal states that requests refused for policy violations are still charged — you pay for rejections
    free tier
    no
    access
    closed api (xai, fal)
    max resolution
    2K
    text rendering
    adequate
    editing
    strong — its best capability
    license
    commercial via api terms

    grok imagine's quality tier ranks seventh on artificial analysis' editing leaderboard, better than its standing for plain generation, so this is more an editing tool than a headline generator.

    its content policy is considerably more permissive than google's or openai's, which for some legitimate use cases — medical, security, certain editorial work — is the difference between a usable tool and a wall of refusals.

    the thing to flag clearly: fal's model page states that when a request is deemed to violate xai's terms, the generation is still charged. paying for refusals is unusual and it changes the economics of any pipeline that hits the policy boundary with any regularity.

    xai also carries meaningful reputational baggage around deepfakes and non-consensual imagery, including litigation and a multistate attorney-general letter. that is a brand-safety question worth asking before you put it in a consumer product.

    pros
    • +#7 on the editing leaderboard, strong for the price
    • +$0.02 on the standard tier is very cheap
    • +permissive policy unblocks legitimate edge cases
    • +fast, with a simple api
    cons
    • you are charged for generations it refuses
    • significant deepfake-related litigation and regulatory attention
    • 2K resolution ceiling
  16. 16

    Midjourney V8.1

    the best house style in the business, still with no api

    68/100

    verdictstill the most beautiful default output anywhere, and still impossible to build a product on — no public api, four years running.

    best for
    art direction and mood work where the image has to look good before it looks correct
    price
    $10 / month basic, up to $120 / month mega
    pricing note
    20% off annual; stealth mode (private generations) requires the $60 pro tier or above
    free tier
    no
    access
    subscription web app, no api
    max resolution
    2K native
    text rendering
    improved, still not a strength
    editing
    region inpaint, style/char refs
    license
    commercial on paid plans

    midjourney's advantage was always taste. give it a lazy prompt and you get intentional lighting, composition and colour where most models give you something technically correct and visually inert. v8 rewrote the codebase and v8.1 followed in april, bringing roughly 5× faster generation, native 2K and better text.

    for concept work, mood boards and anything with a feel brief, that head start is still worth real money.

    and it is still subscription-only with no public api. you cannot put it in a product, a pipeline, or anything automated. in a year when every competitor ships an api on day one, that has moved from quirk to disqualifier for most commercial buyers.

    it also doesn't participate in any public benchmark, so there is no independent measure of where it actually stands. our placement reflects commercial usability as much as output quality — on pure aesthetics it would rank considerably higher. note too that midjourney's site blocks automated access, so the version and pricing details here come from secondary sources and are worth confirming yourself.

    pros
    • +best-looking default output with the least prompt effort
    • +v8.1 brought ~5× speed, native 2K and better text
    • +style and character reference controls remain excellent
    • +large community and a deep corpus of shared styles
    cons
    • no public api — cannot be embedded in a product
    • participates in no benchmark, so no independent measure exists
    • privacy requires the $60/month tier
  17. 17

    FLUX.2 [klein] 4B

    apache 2.0, sub-second, runs on a consumer gpu

    66/100

    verdictgenuinely apache 2.0 and fast enough to feel interactive on a consumer gpu — just make absolutely sure you downloaded the 4b and not the 9b.

    best for
    local and on-device generation where latency and licence freedom matter more than peak quality
    price
    free — open weights
    pricing note
    apache 2.0 — but only the 4B version; the 9B released the same day is non-commercial
    free tier
    yes
    access
    open weights, self-hosted
    max resolution
    1–2K typical
    text rendering
    basic
    editing
    via flux community tooling
    license
    Apache 2.0 (4B only)

    klein 4b runs in about thirteen gigabytes of vram with sub-second generation, which puts real-time local image generation on hardware people actually own. for on-device features, offline tools, or anything where round-tripping to an api is the bottleneck, that matters more than a few elo points.

    it is apache 2.0. no revenue cap, no non-commercial clause, no negotiation.

    here is the trap, and it is the sharpest one in this category. black forest labs released klein 4b and klein 9b on the same day under different licences. the 9b performs better, benchmarks better, and is the one most people reach for — and it ships under the flux non-commercial licence. it is very easy to evaluate the 9b, like it, ship it, and be in breach without ever having read a licence file.

    quality-wise, 4b is a small model and it shows. this is a latency-and-licence pick, not a quality one.

    pros
    • +genuine apache 2.0 — unrestricted commercial use
    • +sub-second generation in ~13gb of vram
    • +viable for on-device and offline products
    • +inherits the flux tooling ecosystem
    cons
    • the better-performing 9b sibling is non-commercial — very easy to get wrong
    • small-model quality, clearly behind the frontier
    • needs local gpu infrastructure
  18. 18

    Luma Uni-1

    reasoning-first generation, quietly competent at editing

    63/100

    verdicta decent reasoning-first model whose editing punches above its overall standing — but photon is long gone and the cheap tier's output isn't commercially usable.

    best for
    teams already using luma for video who want one vendor
    price
    $0.0404 / image at 2048px; uni-1-max $0.10
    pricing note
    reference images add ~$0.003 each, up to nine; the cheapest consumer tier produces watermarked, non-commercial output
    free tier
    yes
    access
    closed api + web app
    max resolution
    2048px
    text rendering
    adequate
    editing
    strong for its tier
    license
    commercial on higher tiers only

    first, a correction to a lot of stale coverage: photon is not luma's current image model. uni-1 replaced it in march, with uni-1.1 following on the api in may. anything recommending luma photon is out of date.

    uni-1 reasons about the prompt before generating, and uni-1 max sits ninth on artificial analysis' editing leaderboard — a better showing than its overall generation ranking, which suggests editing is where the reasoning pays off.

    at $0.0404 per image at 2048px it is priced in the same band as nano banana 2 without matching it on capability, so the main reason to pick it is vendor consolidation if you are already using luma for video.

    watch the consumer tiers: the entry plan produces watermarked output under non-commercial terms. that is a normal restriction, but it catches people who assume a paid plan means commercial rights.

    pros
    • +#9 on the editing leaderboard
    • +reasoning-first architecture, strong instruction following
    • +up to nine reference images
    • +one vendor if you already use luma for video
    cons
    • entry consumer tier output is watermarked and non-commercial
    • priced against nano banana 2 without matching it
    • much stale documentation still refers to photon
  19. 19

    Adobe Firefly Image 5

    mediocre model, unmatched legal cover

    61/100

    verdictit appears in no leaderboard top ten and it doesn't matter — firefly is bought for licensed training data and ip indemnification, and on that it has no competition.

    best for
    agencies and enterprises where training-data provenance is a procurement requirement
    price
    included with Creative Cloud; standalone from ~$10 / month
    pricing note
    generative-credit based; api pricing exists but adobe doesn't publish a clear per-image rate
    free tier
    yes
    access
    subscription + creative cloud + api
    max resolution
    4mp native
    text rendering
    adequate
    editing
    excellent inside photoshop
    license
    licensed data, indemnified

    firefly image 5 does native four-megapixel output and generative fill inside photoshop is still the best implementation of editing-in-context anywhere. that integration, not the model, is the product.

    the actual reason firefly sells is legal. it is trained on adobe stock and licensed content, and adobe offers ip indemnification to enterprise customers. in a year when bytedance's video model drew cease-and-desists from disney and paramount, that story is worth more to a lot of buyers than any amount of image quality.

    as a generator it is unremarkable — competent, conservative, stock-looking, and absent from every leaderboard's top ten. adobe appears to know this, which is why firefly now routes to thirty-plus third-party models including google's, openai's and runway's.

    the generative-credit system remains easy to burn through without noticing, and adobe does not publish clean per-image api pricing.

    pros
    • +licensed training data with enterprise ip indemnification
    • +generative fill inside photoshop is best-in-class in practice
    • +native 4mp output
    • +now aggregates 30+ third-party models in one interface
    cons
    • own model appears in no leaderboard top ten
    • generative credits deplete faster than expected
    • no transparent per-image api pricing
  20. 20

    Stable Diffusion 3.5

    the ecosystem that started it all, twenty-one months without a successor

    48/100

    verdictstill the widest tooling ecosystem in open image models, and now a full generation behind — there is no stable diffusion 4, whatever you have read.

    best for
    existing sd pipelines and techniques that were only ever implemented for sd
    price
    free below $1M annual revenue
    pricing note
    stability community licence; above $1M revenue an enterprise licence is required
    free tier
    yes
    access
    open weights + api
    max resolution
    1–2K typical
    text rendering
    weak by 2026 standards
    editing
    everything, via the ecosystem
    license
    community licence, $1M revenue cap

    stable diffusion 3.5 shipped in october 2024 and remains stability's newest image model. that is twenty-one months without a successor while the rest of the category shipped two or three generations.

    to be explicit, because there is a lot of confident misinformation about this: there is no stable diffusion 4. stability's own model page lists 3.5 as current, and its only 2026 announcements are stable audio 3.0 and brand studio — neither an image model. articles describing sd4 tiers and 4096px output are fabricated.

    what remains genuinely valuable is the ecosystem. comfyui workflows, controlnets, tutorials, and an enormous library of community fine-tunes — if a technique exists, it was implemented for sd first and sometimes only for sd. sdxl in particular is still the right pick when vram is tight or you depend on an sdxl-only lora.

    the community licence is also still generous: free commercial use below a million dollars in annual revenue. as a starting point for a product that might not get big, that is a reasonable place to begin — just don't expect frontier output.

    pros
    • +widest tooling and workflow support of any image model
    • +free commercial use below $1M annual revenue
    • +runs on modest consumer gpus, especially sdxl
    • +enormous library of community fine-tunes and controlnets
    cons
    • no new image model in twenty-one months
    • quality now a full generation behind
    • stability has pivoted to audio and brand tooling

how this ranking was made

hands-on prompting across photoreal, illustration, text-in-image and compositional briefs ('a red mug to the left of a closed laptop'), looking at first generations rather than best-of-ten. this is qualitative — we don't derive a score from it, we use it to sanity-check what the public leaderboards say.

cost per image is the vendor's own published list price for the cheapest path to a usable 1024px result, taken from the vendor's pricing page or the fal model page on the review date. where a model is billed by token or by megapixel we say so rather than inventing a per-image figure.

we read the actual licence file, not the launch blog. in 2026 those disagree often enough that it matters — black forest labs' own flux.2 announcement reads as though [dev] is apache 2.0, and the licence on hugging face is not.

we cross-check our own impressions against the artificial analysis and lmarena image leaderboards, because two independent public vote pools are harder to fool than one reviewer's taste. where our placement disagrees with them, the entry says why.

prices and versions move fast in this category. the date at the top is when we last re-checked every figure on this page.

our general methodology and disclosures →
was this useful?