raw fidelity stopped being the differentiator some time ago. everything in the top half of this list makes a technically clean image. what separates them now is whether the model does what the prompt actually said, whether it can render legible text, whether you can edit one thing without the rest of the frame drifting, and what a thousand images actually cost.
licensing is the axis people keep underweighting, and 2026 made it worse. several of the most-downloaded 'open' models — flux.2 dev, flux.2 klein 9b, ideogram 4 — ship weights under non-commercial licences. teams benchmark them, like them, and ship into breach. the two genuinely permissive frontier-adjacent open models right now are nvidia's cosmos3 and hidream-o1, which is a completely different cast than a year ago.
one structural note on price: routing google's models through an aggregator like fal costs roughly 12–25% more than going direct to the gemini api, and google also offers 50% batch pricing that aggregators don't expose. the convenience of one unified api is real, but so is the markup.
two models are deliberately absent. meta's muse image would rank around third on public leaderboards but is consumer-app-only with no api. alibaba's qwen-image 3.0 shipped on 21 july with no weights, no licence, no paper and no benchmarks — nothing independently testable. we'll rank both when you can actually buy them.
advertisement
1
GPT Image 2
the only model that leads both generation and editing
95/100
verdictthe best model in the category on every measure that can be measured — and priced so that you should still draft on something else.
best for
typography, diagrams, and any prompt with a dozen constraints that all have to survive
price
$0.006 low / $0.053 medium / $0.211 high, at 1024×1024
pricing note
billed per image by quality tier; a 35× spread between low and high makes forecasting genuinely hard
free tier
no
access
closed api + chatgpt
max resolution
3840px max edge, 3:1 aspect limit
text rendering
best in class
editing
inpaint, outpaint, multi-reference
license
commercial via api terms
gpt image 2 is the only model that sits at number one on artificial analysis for both text-to-image and editing, and number one on lmarena with the largest vote pool of any entrant. that combination has not happened before; usually the aesthetic leader and the instruction-following leader are different products.
its signature strength is fine typography. multi-line legible text, correct spelling, sensible kerning, text that sits properly inside a shape — for infographics, packaging mockups and social assets with real copy, it is routinely the only model that gets it right on the first generation.
the reason it isn't an automatic default is the price curve. high quality at 1024×1024 is $0.211 per image, roughly three times nano banana 2 and six times its lite sibling. the low tier is $0.006, which is competitive, but the gap between them is 35× — so a pipeline that quietly defaults to high will produce a bill nobody forecast.
one editing caveat straight from openai's own docs: masking is prompt-based rather than pixel-exact, so it may not follow the precise shape you painted. it is still the best inpainting available; it just isn't the surgical tool the word 'mask' implies.
pros
+#1 on artificial analysis for generation and editing, and #1 on lmarena
+best fine typography and multi-line text of any model
+holds long prompts with many simultaneous constraints
+quality tiers let you draft cheap and pay only for keepers
cons
−$0.211 per high-quality image is the priciest mainstream option
−35× price spread between tiers makes cost forecasting hard
the independent lab that reasons about layout before it draws
90/100
verdictsecond on both public leaderboards and the only model here with genuine 16-megapixel native output — the strongest thing built outside a hyperscaler.
best for
dense compositions, posters, and anything that needs to stay coherent at very large sizes
price
~$0.04 / image on fal's reve endpoint
pricing note
reve has not published 2.1-specific pricing — the $0.04 figure is fal's unversioned reve endpoint, so confirm before budgeting
free tier
yes
access
closed api + web app
max resolution
4096×4096 native (16mp)
text rendering
strong, incl. non-latin
editing
instruction edit + 8-image remix
license
commercial via api terms
reve's architecture reasons about structure and hierarchy before rendering, rather than denoising its way to a composition and hoping the layout survives. in practice that shows up exactly where you would expect: dense posters, editorial layouts, and prompts with many elements that all need their own space stay coherent instead of collapsing into mush.
it renders at 4096×4096 natively — sixteen megapixels, not an upscale. nothing else on this list does that, and for print work or anything that gets cropped hard afterwards it removes a whole post-processing step.
it sits second on both artificial analysis and lmarena, which is a genuinely surprising result for a company this size. the caveat worth knowing is that its lmarena vote count is a fraction of gpt image 2's, so the ranking is real but less settled — expect it to move more than the top spot does.
editing is covered by two modes: instruction-based edit on a single image, and remix across up to eight references. the missing piece is pricing transparency — reve does not publish a per-image rate for 2.1 anywhere we could find, which is an odd gap for a product this good.
pros
+#2 on both artificial analysis and lmarena
+4096×4096 native output, the highest here by a wide margin
google's gemini 3.1 flash image — the best default for production work
89/100
verdictnot the top of any leaderboard, and still the model we'd default to — the balance of price, speed, resolution and character consistency is better than anything above it.
best for
high-volume commercial pipelines that need consistent characters and predictable cost
price
$0.067 / image at 1K direct from google
pricing note
$0.045 at 0.5K, $0.101 at 2K, $0.151 at 4K; fal charges ~19% more; google offers 50% batch pricing that aggregators don't expose
free tier
yes
access
closed api (gemini, vertex) + apps
max resolution
512px to 4K
text rendering
strong, incl. in-image translation
editing
excellent targeted editing
license
commercial; synthid watermarked
the naming is a mess, so to be clear: nano banana 2 is gemini 3.1 flash image, released in february. it is not a successor to nano banana pro in the quality sense — it is the fast line catching up, which it comprehensively has.
its standout capability is subject consistency: up to five characters and fourteen objects held stable across generations. for anyone producing a series — product shots, a character across a storyboard, a campaign that has to look like one campaign — that is worth more than a few elo points, and it is class-leading.
it ranks third on artificial analysis' editing board, ahead of the more expensive nano banana pro, and does 4K, web grounding and in-image translation. at $0.067 per image at 1K it costs less than a third of gpt image 2 at high quality.
where it loses is raw aesthetic preference — head to head, voters pick gpt image 2 by a wide margin. if a single hero image has to be the best possible, this isn't it. if you need ten thousand good ones, it is.
pros
+best price/quality/speed balance on the market
+five-character, fourteen-object consistency is class-leading
+#3 on the editing leaderboard, ahead of nano banana pro
+4K output, web grounding, and 50% batch pricing direct from google
cons
−clearly behind gpt image 2 on aesthetic preference
−the nano banana naming makes version choice genuinely confusing
−~19% markup if you route through fal instead of google
the budget tier that outscores the flagship it's cheaper than
87/100
verdictscores above full nano banana 2 on artificial analysis at half the price and nearly three times the speed — the value story of 2026.
best for
high-volume generation where 1024px is enough and latency matters
price
$0.034 / image at 1K direct from google
pricing note
on fal it is billed by token at $37.50/1M image-out with fixed 1024×1024 output — roughly 25% more than going direct
free tier
yes
access
closed api (gemini, vertex)
max resolution
1024×1024 via fal
text rendering
strong for the price
editing
supported, flash-tier quality
license
commercial; synthid watermarked
this is the entry that breaks the usual assumption. released at the end of june, nano banana 2 lite posts a higher artificial analysis elo than the full nano banana 2 it sits beneath in google's own lineup, at $0.034 per image against $0.067, and generates in about four seconds — roughly 2.7× faster.
cheap tiers being strictly worse used to be a safe rule. it isn't any more, and if you are generating at volume you should benchmark this before paying for anything above it.
the constraint is resolution. through fal it is pinned to 1024×1024, so anything needing 2K or 4K has to move up the family. it also inherits the flash line's aesthetic — competent and slightly literal rather than striking.
the billing model on fal is token-based rather than per-image, which makes cost modelling more annoying than it should be. going direct to google is both cheaper and easier to reason about.
pros
+outscores full nano banana 2 on artificial analysis
+half the price of nano banana 2, ~2.7× faster
+$0.034/image is frontier-adjacent quality at budget cost
+same google reliability, sdks and batch discount
cons
−fixed 1024×1024 output through fal — no 2K or 4K
−token-based billing on aggregators is awkward to forecast
microsoft's own model, strong scores, awkward to buy
84/100
verdictthird on artificial analysis and genuinely excellent at text and stylised art — held back almost entirely by how hard it is to buy if you aren't already on azure.
best for
teams already inside azure or microsoft foundry
price
token-billed: $47 / 1M image-output tokens
pricing note
microsoft publishes token rates only, no per-image figure; the Flash variant is $33/1M image-out
free tier
no
access
closed api (foundry, openrouter)
max resolution
not published
text rendering
excellent
editing
control-with-preservation editing
license
commercial via api terms
released in june, mai-image-2.5 is microsoft's own image model rather than a rebadged openai one, and it is good: third on artificial analysis text-to-image, fourth on editing, sixth on lmarena. the jump over the previous generation was largest in text rendering and in cartoon, anime and fantasy styles.
its editing approach — microsoft calls it control with preservation — is aimed squarely at the failure mode where a targeted change quietly redraws the rest of the frame, and it handles that well.
distribution is the problem. it is available through microsoft foundry, the mai playground and openrouter, and it is shipping inside powerpoint and onedrive, but there is no clean consumer-facing api story and microsoft publishes token rates rather than a per-image price. working out what a thousand images costs takes a spreadsheet.
if you are already an azure shop, this is a strong and under-discussed option. if you aren't, the friction is real and the models above it are easier to adopt.
pros
+#3 on artificial analysis text-to-image, #4 on editing
+large generational jump in text rendering and stylised art
+editing preserves the rest of the frame well
+already embedded in powerpoint and onedrive
cons
−token-only pricing with no published per-image cost
google's gemini 3 pro image — the accuracy specialist, now overpriced
83/100
verdictstill the best model for grounded, factually accurate imagery — but its own cheaper stablemate now beats it at editing, which makes the price hard to defend.
best for
infographics and anything where factual accuracy in the image matters more than cost
price
$0.134 / image at 1K–2K, $0.24 at 4K
pricing note
fal charges $0.15 and $0.30, plus $0.015 when web search is used
free tier
yes
access
closed api (gemini, vertex) + apps
max resolution
4K
text rendering
strong
editing
strong, #5 on editing board
license
commercial; synthid watermarked
nano banana pro is gemini 3 pro image, and its distinguishing trick is that it can search the web mid-generation. for infographics, diagrams with real data, or anything where the picture makes a factual claim, that grounding produces results the others simply can't.
it was google's quality flagship from november until february, and google's own guidance still points here for high-fidelity work requiring maximum factual accuracy.
the problem is internal competition. nano banana 2 ranks above it on the editing leaderboard at half the price, and nano banana 2 lite outscores it on price-performance by a wider margin still. unless you specifically need the web grounding or 4K at this tier, the cheaper siblings are the better buy.
it is also the slow one in the family. for interactive or high-volume work the latency is noticeable.
pros
+web grounding produces genuinely factual imagery
+best-in-family world knowledge for infographics
+native 4K output
+google reliability, sdks and batch pricing
cons
−beaten on editing by nano banana 2 at half the price
−slowest of the nano banana family
−$0.24 at 4K is expensive for what it now delivers
the best non-latin text rendering available anywhere
82/100
verdictif your text isn't in english, this is the model — ten-plus languages rendered natively, and nothing else comes close on cjk.
best for
multilingual design work, especially cjk and other non-latin scripts
price
$0.0675 / image up to 1536×1536
pricing note
$0.135 from there to 2048×2048; available via volcano engine, byteplus and fal
free tier
no
access
closed api (volcano, byteplus, fal)
max resolution
2048×2048, ~3K long edge
text rendering
best for non-latin scripts
editing
layered + precise interactive edit
license
commercial via api terms
bytedance released seedream 5.0 pro two weeks ago, and its differentiator is unambiguous: native multilingual text rendering across more than ten languages. western models fall apart on chinese, japanese and korean characters in a way that is obvious to anyone who reads them. this one doesn't.
it is also built for information density — charts, dense infographics, layouts with a lot of small type — and supports multi-layer separation and precise interactive editing, which makes it unusually practical for design handoff rather than just generation.
at $0.0675 per image up to 1536², it is priced competitively against nano banana 2 while doing something none of the models above it can.
two caveats. it is brand new and does not yet appear in either leaderboard's top ten, so our placement leans on capability rather than public voting. and for western enterprises there is procurement friction around a bytedance model that is worth raising internally before it becomes a surprise.
pros
+best non-latin and cjk text rendering by a clear margin
+ten-plus languages rendered natively
+strong on dense information graphics and layered output
+competitively priced at $0.0675
cons
−absent from both leaderboard top tens despite the capability
−bytedance procurement friction for some western buyers
−2048×2048 ceiling, lower than its own lite sibling
verdictthe best open model you can legally use in a commercial product — and it beats models seven times its size.
best for
self-hosted commercial products that need weights you can actually ship on
price
free — open weights
pricing note
MIT licence, no revenue cap, no commercial restriction; you pay only for gpus
free tier
yes
access
open weights, self-hosted
max resolution
model-dependent, typically 1–2K
text rendering
good
editing
via community tooling
license
MIT — fully commercial
hidream-o1 is pixel-native: no vae, no separate text encoder, and it reasons before it draws. at 8 billion parameters it trades punches with flux.2 dev at 32 billion, and its 1.5 revision sits fourth on artificial analysis.
the reason it ranks this high is the licence. it is MIT. no revenue cap, no non-commercial clause, no 'contact us above 100 million users'. in a year where most of the open-weight headline releases turned out to be non-commercial, that is the differentiator that actually matters.
self-hosting still means gpu capacity, ops, and no vendor to call at 2am. this is the right pick when you need control, data residency, or per-image economics that only work if inference is yours — not when you want the shortest path to an image.
ecosystem is the weak point. the lora and controlnet tooling that grew up around flux does not exist here in the same depth, so expect more integration work.
pros
+MIT licence — unrestricted commercial use
+competitive with models 4× its parameter count
+#4 on artificial analysis for the 1.5 revision
+pixel-native architecture runs without a separate text encoder
ten reference images and the deepest tooling ecosystem in the category
79/100
verdictno longer leads any leaderboard, but nothing else takes ten reference images — for product and brand work that single capability still wins.
best for
brand and product consistency across large volumes of images
price
$0.03 first megapixel + $0.015 per extra megapixel
pricing note
fal matches black forest labs' direct pricing here — no aggregator markup, unlike the google models
free tier
no
access
closed api (bfl, fal)
max resolution
4mp editing
text rendering
good
editing
10-reference composition, 4mp
license
commercial via api; weights are not
flux.2 pro accepts up to ten reference images and lets you address them individually in the prompt. for keeping a product, a face or a brand style consistent across hundreds of generations, that is a different class of control from a single style reference, and it is the main reason this model is still in the top ten.
it edits at four megapixels, handles photorealism well, and its pricing is honest — $0.03 for the first megapixel and $0.015 for each additional one, matching black forest labs' direct rate rather than marking it up.
it has, however, fallen out of the leaderboard top ten entirely. a year ago flux was the answer; today it is a specialist with an unusually deep toolchain around it.
the family is also a licensing minefield, and this is the entry to read carefully. the api models are fine — you are buying hosted inference. the downloadable ones are not: flux.2 dev and flux.2 klein 9b both ship under the flux non-commercial licence despite launch messaging that reads otherwise. generated outputs can be used commercially; the weights cannot.
pros
+up to ten reference images, individually addressable
+best-in-class product and brand consistency
+4mp editing and strong photorealism
+no aggregator markup — fal matches bfl direct pricing
cons
−no longer in the leaderboard top ten
−the wider flux family's licensing traps catch a lot of teams
verdictthe top open-weight model on artificial analysis, released with datasets and training recipes — if provenance matters to you, nothing else is close.
best for
research, fine-tuning, and teams who need to see what the model was trained on
price
free — open weights
pricing note
OpenMDW-1.1 licence with commercial rights; checkpoints, six datasets and training recipes all published
free tier
yes
access
open weights, self-hosted
max resolution
model-dependent
text rendering
adequate
editing
via community tooling
license
OpenMDW-1.1 — commercial ok
cosmos3 is the highest-ranked open-weight model on artificial analysis, and nvidia released it properly: checkpoints, six datasets, and the training recipes. most 'open' releases in 2026 mean a weights file and a licence that forbids using it. this one means what the word used to mean.
the openmdw-1.1 licence grants commercial rights, which together with hidream's MIT makes these two the only genuinely shippable open models near the frontier.
it is also agentic — it plans and revises rather than denoising straight from the prompt — which is where the whole category moved this year.
practically, it is a heavier lift than hidream: bigger, more infrastructure, and aimed at teams who want to fine-tune rather than teams who want to call an api. if you need auditable training data for a regulated client, the dataset release is worth the effort on its own.
pros
+highest-ranked open-weight model on artificial analysis
+openmdw-1.1 licence permits commercial use
+datasets and training recipes published, not just weights
+agentic planning rather than straight denoising
cons
−heavier infrastructure requirement than hidream-o1
still the typography specialist, with a licence you must read
76/100
verdictstill the best model for stylised lettering, and the cheapest good option per megapixel — but the downloadable weights are non-commercial, whatever the press coverage says.
best for
posters, packaging, signage — anything where the words are the design
a 2048×2048 image runs $0.021 to $0.07 depending on tier
free tier
yes
access
closed api + web + weights
max resolution
2K native
text rendering
best for stylised type
editing
inpaint, remix, layout boxes
license
weights non-commercial; api commercial
ideogram built its identity on rendering text correctly and it still leads on typography specifically: dense small type, multi-word headlines, lettering that follows a curve or sits properly inside a shape. it is also number one open-weight model on designarena.
version 4 added bounding-box layout control and a structured json prompting interface, which is genuinely differentiated — if you are generating hundreds of assets that each carry a headline in a fixed position, that turns a creative tool into a pipeline component.
at $0.00525 per megapixel on the turbo tier it is also among the cheapest credible models here.
now the important part. ideogram 4 is widely reported as open weights under a commercial licence. the model cards on hugging face say otherwise — both the fp8 and nf4 releases carry an ideogram 4 non-commercial agreement, and fal's own derivative states that it inherits that non-commercial term. treat the weights as non-commercial and the hosted api as the commercial path, and get written confirmation from ideogram before you plan otherwise.
pros
+category leader for stylised typography and lettering
+bounding-box layout control and json prompting
+cheapest quality tier per megapixel on this list
+generous free tier for evaluation
cons
−downloadable weights are non-commercial despite press claims
verdictgenuine editable svg output with no real competitor — for design systems it isn't the best model, it's the only one.
best for
icon sets, logos, and design-system work that has to stay editable
price
~$0.035 / image (V4.1); V4 Pro ~$0.25
pricing note
recraft's own pricing page doesn't render figures — these come from aggregators, so confirm before budgeting
free tier
yes
access
closed api + web app
max resolution
vector — resolution independent
text rendering
good in vector layouts
editing
vector editing + style sets
license
commercial on paid plans
recraft generates actual vector graphics: editable svg, not a raster image auto-traced afterwards. for icon sets, logo exploration and spot illustration that go straight into figma and stay editable, that is a category difference rather than a feature.
brand style sets are the other strong idea — train a style from references and every subsequent generation stays on-brand, which is the consistency problem that makes most image models unusable for production design systems.
v4.1 utility pro sits tenth on artificial analysis, which is a respectable showing for a tool this specialised.
photoreal work is not its strength and it doesn't pretend otherwise. it is a design tool that happens to use generative models and should be judged that way. pricing is also frustratingly hard to pin down — recraft's own pricing page didn't render figures for us, so the numbers here come from aggregators.
pros
+true editable svg output — effectively unique
+brand style sets give real cross-asset consistency
+#10 on artificial analysis despite the narrow focus
+strong at icons, infographics and spot illustration
the best aesthetics-per-dollar, built around style control
74/100
verdictthe cheapest route to a genuinely good-looking image, with the most thoughtful style controls in the category.
best for
art direction, moodboarding, and iterating on a look rather than a fact
price
$0.008 / megapixel
pricing note
one of the cheapest quality models available; krea 2 turbo generates in roughly two seconds
free tier
yes
access
closed api + web app
max resolution
megapixel-billed, flexible
text rendering
weak
editing
style refs, loras, sliders
license
commercial on paid plans
krea's whole thesis is style control. moodboards, multiple style references with adjustable influence weights, lora support, and generative sliders for intensity, complexity and movement. it is built for people who know what they want it to look like and are tired of describing it in adjectives.
krea 2 turbo generates in about two seconds at $0.008 per megapixel, which makes it viable for the kind of rapid iteration that is prohibitively expensive on gpt image 2.
it is not a text-rendering or factual-accuracy model, and it doesn't try to be. put a headline in the prompt and you will be disappointed.
one thing to be careful with: krea markets itself as sixth globally on artificial analysis and first among independent labs. we could not corroborate the sixth placement on the leaderboard we pulled, where it sits outside the top ten. the model is good and the price is excellent — the self-reported ranking is the part we'd discount.
pros
+$0.008/mp is among the cheapest quality options anywhere
+best style-reference and moodboard controls in the category
+generative sliders for intensity, complexity and movement
+roughly two-second generation
cons
−weak text rendering — not a typography tool
−self-reported leaderboard ranking we couldn't verify
higher resolution than the pro model, at half the price
72/100
verdictan odd product — cheaper and higher-resolution than the pro tier above it, and worth benchmarking before you assume pro is the upgrade.
best for
cheap high-resolution generation, especially with multilingual text
price
$0.035 / image
pricing note
outputs up to 3072×3072 — counterintuitively higher than seedream 5.0 pro's 2048 ceiling
free tier
no
access
closed api (volcano, byteplus, fal)
max resolution
3072×3072
text rendering
strong, multilingual
editing
basic relative to pro
license
commercial via api terms
seedream 5.0 lite outputs at up to 3072×3072, which is larger than seedream 5.0 pro's 2048 ceiling, at roughly half the price. that is not how product tiers normally work, and it is worth knowing before you default to the pro model.
it carries the same architecture combining deep thinking, web search and generation, and inherits the family's strong multilingual text rendering.
what you give up is the pro model's layer separation, its precise interactive editing, and some quality at the top end. for a lot of production work that trade is fine.
as with the pro model, it is new enough that public leaderboards haven't placed it, and bytedance procurement questions apply equally.
pros
+3072×3072 output — higher than the pro tier
+$0.035 per image is genuinely cheap for the resolution
+inherits strong multilingual text rendering
+unified thinking, search and generation architecture
cons
−no layer separation or precise interactive editing
−unplaced on public leaderboards
−same bytedance procurement friction as the pro model
good editing, permissive content policy, and a billing practice you should know about
69/100
verdicta solid editing model with a notably permissive content policy — and the only one here that bills you for generations it refuses to produce.
best for
editing work and subjects other models refuse
price
$0.02 standard, $0.05 at 1K / $0.07 at 2K on the quality tier
pricing note
fal states that requests refused for policy violations are still charged — you pay for rejections
free tier
no
access
closed api (xai, fal)
max resolution
2K
text rendering
adequate
editing
strong — its best capability
license
commercial via api terms
grok imagine's quality tier ranks seventh on artificial analysis' editing leaderboard, better than its standing for plain generation, so this is more an editing tool than a headline generator.
its content policy is considerably more permissive than google's or openai's, which for some legitimate use cases — medical, security, certain editorial work — is the difference between a usable tool and a wall of refusals.
the thing to flag clearly: fal's model page states that when a request is deemed to violate xai's terms, the generation is still charged. paying for refusals is unusual and it changes the economics of any pipeline that hits the policy boundary with any regularity.
xai also carries meaningful reputational baggage around deepfakes and non-consensual imagery, including litigation and a multistate attorney-general letter. that is a brand-safety question worth asking before you put it in a consumer product.
pros
+#7 on the editing leaderboard, strong for the price
+$0.02 on the standard tier is very cheap
+permissive policy unblocks legitimate edge cases
+fast, with a simple api
cons
−you are charged for generations it refuses
−significant deepfake-related litigation and regulatory attention
the best house style in the business, still with no api
68/100
verdictstill the most beautiful default output anywhere, and still impossible to build a product on — no public api, four years running.
best for
art direction and mood work where the image has to look good before it looks correct
price
$10 / month basic, up to $120 / month mega
pricing note
20% off annual; stealth mode (private generations) requires the $60 pro tier or above
free tier
no
access
subscription web app, no api
max resolution
2K native
text rendering
improved, still not a strength
editing
region inpaint, style/char refs
license
commercial on paid plans
midjourney's advantage was always taste. give it a lazy prompt and you get intentional lighting, composition and colour where most models give you something technically correct and visually inert. v8 rewrote the codebase and v8.1 followed in april, bringing roughly 5× faster generation, native 2K and better text.
for concept work, mood boards and anything with a feel brief, that head start is still worth real money.
and it is still subscription-only with no public api. you cannot put it in a product, a pipeline, or anything automated. in a year when every competitor ships an api on day one, that has moved from quirk to disqualifier for most commercial buyers.
it also doesn't participate in any public benchmark, so there is no independent measure of where it actually stands. our placement reflects commercial usability as much as output quality — on pure aesthetics it would rank considerably higher. note too that midjourney's site blocks automated access, so the version and pricing details here come from secondary sources and are worth confirming yourself.
pros
+best-looking default output with the least prompt effort
+v8.1 brought ~5× speed, native 2K and better text
+style and character reference controls remain excellent
+large community and a deep corpus of shared styles
cons
−no public api — cannot be embedded in a product
−participates in no benchmark, so no independent measure exists
verdictgenuinely apache 2.0 and fast enough to feel interactive on a consumer gpu — just make absolutely sure you downloaded the 4b and not the 9b.
best for
local and on-device generation where latency and licence freedom matter more than peak quality
price
free — open weights
pricing note
apache 2.0 — but only the 4B version; the 9B released the same day is non-commercial
free tier
yes
access
open weights, self-hosted
max resolution
1–2K typical
text rendering
basic
editing
via flux community tooling
license
Apache 2.0 (4B only)
klein 4b runs in about thirteen gigabytes of vram with sub-second generation, which puts real-time local image generation on hardware people actually own. for on-device features, offline tools, or anything where round-tripping to an api is the bottleneck, that matters more than a few elo points.
it is apache 2.0. no revenue cap, no non-commercial clause, no negotiation.
here is the trap, and it is the sharpest one in this category. black forest labs released klein 4b and klein 9b on the same day under different licences. the 9b performs better, benchmarks better, and is the one most people reach for — and it ships under the flux non-commercial licence. it is very easy to evaluate the 9b, like it, ship it, and be in breach without ever having read a licence file.
quality-wise, 4b is a small model and it shows. this is a latency-and-licence pick, not a quality one.
pros
+genuine apache 2.0 — unrestricted commercial use
+sub-second generation in ~13gb of vram
+viable for on-device and offline products
+inherits the flux tooling ecosystem
cons
−the better-performing 9b sibling is non-commercial — very easy to get wrong
reasoning-first generation, quietly competent at editing
63/100
verdicta decent reasoning-first model whose editing punches above its overall standing — but photon is long gone and the cheap tier's output isn't commercially usable.
best for
teams already using luma for video who want one vendor
price
$0.0404 / image at 2048px; uni-1-max $0.10
pricing note
reference images add ~$0.003 each, up to nine; the cheapest consumer tier produces watermarked, non-commercial output
free tier
yes
access
closed api + web app
max resolution
2048px
text rendering
adequate
editing
strong for its tier
license
commercial on higher tiers only
first, a correction to a lot of stale coverage: photon is not luma's current image model. uni-1 replaced it in march, with uni-1.1 following on the api in may. anything recommending luma photon is out of date.
uni-1 reasons about the prompt before generating, and uni-1 max sits ninth on artificial analysis' editing leaderboard — a better showing than its overall generation ranking, which suggests editing is where the reasoning pays off.
at $0.0404 per image at 2048px it is priced in the same band as nano banana 2 without matching it on capability, so the main reason to pick it is vendor consolidation if you are already using luma for video.
watch the consumer tiers: the entry plan produces watermarked output under non-commercial terms. that is a normal restriction, but it catches people who assume a paid plan means commercial rights.
pros
+#9 on the editing leaderboard
+reasoning-first architecture, strong instruction following
+up to nine reference images
+one vendor if you already use luma for video
cons
−entry consumer tier output is watermarked and non-commercial
verdictit appears in no leaderboard top ten and it doesn't matter — firefly is bought for licensed training data and ip indemnification, and on that it has no competition.
best for
agencies and enterprises where training-data provenance is a procurement requirement
price
included with Creative Cloud; standalone from ~$10 / month
pricing note
generative-credit based; api pricing exists but adobe doesn't publish a clear per-image rate
free tier
yes
access
subscription + creative cloud + api
max resolution
4mp native
text rendering
adequate
editing
excellent inside photoshop
license
licensed data, indemnified
firefly image 5 does native four-megapixel output and generative fill inside photoshop is still the best implementation of editing-in-context anywhere. that integration, not the model, is the product.
the actual reason firefly sells is legal. it is trained on adobe stock and licensed content, and adobe offers ip indemnification to enterprise customers. in a year when bytedance's video model drew cease-and-desists from disney and paramount, that story is worth more to a lot of buyers than any amount of image quality.
as a generator it is unremarkable — competent, conservative, stock-looking, and absent from every leaderboard's top ten. adobe appears to know this, which is why firefly now routes to thirty-plus third-party models including google's, openai's and runway's.
the generative-credit system remains easy to burn through without noticing, and adobe does not publish clean per-image api pricing.
pros
+licensed training data with enterprise ip indemnification
+generative fill inside photoshop is best-in-class in practice
+native 4mp output
+now aggregates 30+ third-party models in one interface
the ecosystem that started it all, twenty-one months without a successor
48/100
verdictstill the widest tooling ecosystem in open image models, and now a full generation behind — there is no stable diffusion 4, whatever you have read.
best for
existing sd pipelines and techniques that were only ever implemented for sd
price
free below $1M annual revenue
pricing note
stability community licence; above $1M revenue an enterprise licence is required
free tier
yes
access
open weights + api
max resolution
1–2K typical
text rendering
weak by 2026 standards
editing
everything, via the ecosystem
license
community licence, $1M revenue cap
stable diffusion 3.5 shipped in october 2024 and remains stability's newest image model. that is twenty-one months without a successor while the rest of the category shipped two or three generations.
to be explicit, because there is a lot of confident misinformation about this: there is no stable diffusion 4. stability's own model page lists 3.5 as current, and its only 2026 announcements are stable audio 3.0 and brand studio — neither an image model. articles describing sd4 tiers and 4096px output are fabricated.
what remains genuinely valuable is the ecosystem. comfyui workflows, controlnets, tutorials, and an enormous library of community fine-tunes — if a technique exists, it was implemented for sd first and sometimes only for sd. sdxl in particular is still the right pick when vram is tight or you depend on an sdxl-only lora.
the community licence is also still generous: free commercial use below a million dollars in annual revenue. as a starting point for a product that might not get big, that is a reasonable place to begin — just don't expect frontier output.
pros
+widest tooling and workflow support of any image model
+free commercial use below $1M annual revenue
+runs on modest consumer gpus, especially sdxl
+enormous library of community fine-tunes and controlnets
hands-on prompting across photoreal, illustration, text-in-image and compositional briefs ('a red mug to the left of a closed laptop'), looking at first generations rather than best-of-ten. this is qualitative — we don't derive a score from it, we use it to sanity-check what the public leaderboards say.
cost per image is the vendor's own published list price for the cheapest path to a usable 1024px result, taken from the vendor's pricing page or the fal model page on the review date. where a model is billed by token or by megapixel we say so rather than inventing a per-image figure.
we read the actual licence file, not the launch blog. in 2026 those disagree often enough that it matters — black forest labs' own flux.2 announcement reads as though [dev] is apache 2.0, and the licence on hugging face is not.
we cross-check our own impressions against the artificial analysis and lmarena image leaderboards, because two independent public vote pools are harder to fool than one reviewer's taste. where our placement disagrees with them, the entry says why.
prices and versions move fast in this category. the date at the top is when we last re-checked every figure on this page.