verifier.org

best 7 generative media api platforms

ranked on what one flux.1 [dev] image actually costs, whether you can know the price before you run, and who bills you for work you cannot use.

last reviewed 21 aug 2026 · 7 tools tested ·list curated by Onur Ozcanxin

the short version
best overallfalteams shipping an image or video feature who want the widest model choice and the least thinking91/100runner-upWiroagents and pipelines that need to check the cost of a job before committing to it84/100best free optionRunwarehigh-volume image generation where you can benchmark your own model choice first75/100

these platforms host other people's models. you send a prompt to one api and get an image, a video or a voice back, without renting a gpu or waiting on a cold container. the pitch is that you skip the ops and pay per output, and for most teams shipping a generative feature that pitch is correct.

so we picked the most boring possible test: what does one 1024x1024 image from flux.1 [dev] cost here? almost every platform in this set carries that model, which makes it the only figure that puts them side by side. two of the seven print a directly comparable per-image price. fal charges $0.025 per megapixel and Replicate $0.025 per image, and after that the comparison falls apart — Segmind bills the same model in gpu-seconds, Novita does not carry the base model at all, and Runware, whose entire marketing claim is a $0.0006 price floor, does not surface a per-model figure anywhere public.

that is the real split in this category, and it is not about speed. some platforms bill per output, so the cost of a job is knowable before you run it. others bill per gpu-second, so the cost depends on how long the gpu happened to take, and you find out afterwards. the second kind is cheaper when a model is fast and considerably more expensive when it is not, and no amount of dashboard design makes an agent loop safe to point at it.

one absence worth explaining. Sieve was on the shortlist and is not here: sievedata.com now redirects to sieve.ai, which describes itself as a multimodal data lab selling training data to labs, and its pricing page 404s. it is no longer a generative media api. Baseten and Together were cut deliberately too — both lean toward serving text models, which is a different buying decision, covered in llm inference providers. the model vendors' own apis are not here either, because those belong on the model rankings rather than the platform ones.

advertisement
  1. 1

    fal

    the default, and the only one with a media catalogue this deep

    91/100

    verdictthe most complete media catalogue with per-output pricing printed on every model page — you pay a visible aggregator markup for it, and your credits expire.

    best for
    teams shipping an image or video feature who want the widest model choice and the least thinking
    price
    $0.025 / megapixel for flux.1 [dev]
    pricing note
    per-output pricing on model apis; custom deployments are billed per gpu-hour, with h100 listed at $1.89/hr. purchased credits expire 365 days from purchase and promotional credits in 90
    free tier
    no
    flux.1 [dev]
    $0.025 / megapixel
    billing
    per output; per gpu-hour for deployments
    catalogue
    1,000+ media models
    llms too
    no — media only
    mcp server
    none published

    fal carries over a thousand image, video, audio and 3d models, and prices each one on its own page in the unit that model is actually billed in. that sounds mundane until you compare it with the rest of this list, where the same question takes three page loads and often has no answer. per megapixel for images, per second for video, printed where you are already standing.

    the catalogue is the reason to be here. when a new model lands, fal usually has it within days, and our image and video rankings lean on fal model pages for pricing more than any other source because those pages are the ones that exist. for a team that wants to try seedream against nano banana against flux this afternoon, nothing else is close.

    the markup is real and we have measured it elsewhere on this site: routing google's models through fal costs roughly 12 to 25 percent more than going direct to the gemini api, and google's 50 percent batch pricing is not exposed through aggregators at all. on black forest labs models fal matches direct pricing exactly, so the markup is a per-vendor question rather than a flat tax, but on the google family it is the largest single line item in a heavy pipeline.

    two things to know before you commit budget. fal does not serve text models, so if you want one api across generation and reasoning this is not it. and the billing terms have edges: server errors are never charged, but a 422 client error can still be billed if a runner spent gpu time before the error surfaced, and on at least one model page — Grok Imagine — fal states that generations refused for policy violations are charged anyway. credits also expire, at 365 days for purchased and 90 for promotional, which is a term most of this category does not impose.

    pros
    • +over a thousand media models, and new releases land fast
    • +per-output price printed on every individual model page
    • +matches black forest labs direct pricing with no markup
    • +the de facto reference other platforms are compared against
    cons
    • roughly 12-25% markup on google models versus going direct
    • media only — no text or reasoning models
    • purchased credits expire after 365 days, promotional after 90
    • 422 errors can still be billed when gpu time was spent, and one model page states policy refusals are charged
  2. 2

    Wiro

    the only one that tells you what a job costs before you run it

    84/100

    verdictthe best alternative to fal if billing mechanics matter more to you than catalogue size — it prices every job up front and bills only completed runs, on a catalogue a third the size with almost no community.

    best for
    agents and pipelines that need to check the cost of a job before committing to it
    price
    $0.006 / request for flux-2-dev
    pricing note
    flat per-request on wiro's own optimised endpoints regardless of resolution; community checkpoints are gpu-second billed instead, with flux-1-dev estimated at about $0.031 per run. no public pricing page — wiro.ai/pricing 404s and prices live per model and in the api
    free tier
    no
    flux.1 [dev]
    $0.006 / request (flux-2-dev)
    billing
    six methods, declared per model
    catalogue
    538 models
    llms too
    yes — 93 llm and chat
    mcp server
    yes — mcp.wiro.ai/v1

    wiro's distinguishing feature is not the model list, it is that the price is a queryable field. every model declares one of six billing methods — per request, per second, per token, per pixel, per second of input audio, or per realtime turn — and the api returns dynamicprice and approximatelycost before you submit anything. nothing else in this category lets you ask what a job will cost and get a number back. for an agent that decides on its own how many generations to run, that is the difference between a budget and a surprise.

    the billing terms are the cleanest here as well. wiro states that you are billed for successfully completed model runs, that server errors incur no charge, and that tasks cancelled before processing completes are not billed. fal also refuses to bill server errors, so this is a margin rather than a chasm, but the margins go wiro's way: cancelled work is explicitly free, and we found no credit expiry of the kind fal applies at 365 days.

    the catalogue is 538 models — 167 video, 105 image, 83 image editing, 93 llm and chat, 29 audio, 17 music, 10 realtime and 4 3d. that is a third of fal's media count, and Replicate's community library is larger again. what wiro has instead is coverage across generation and text in one api, plus a realtime conversation tier, and a genuine mcp server at mcp.wiro.ai/v1 that we rank separately and that makes the whole catalogue drivable by an agent without writing a client.

    two pricing worlds live on one platform and you need to know which you are in. wiro's own tuned endpoints are flat and cheap — flux-2-dev is $0.006 a request whatever resolution you ask for, the lowest figure in this ranking. the community checkpoints are gpu-second billed with only an estimate up front, so the same family's flux-1-dev lands near $0.031 a run. and then there is the thing nobody else does: your concurrency is your wallet. below a $250 balance you get concurrent tasks equal to 10 percent of your balance, minimum one, so ten parallel jobs means holding $100. above $250 the cap disappears entirely. it is a defensible anti-abuse design and it is still a prepayment requirement dressed as a rate limit.

    the honest weakness is adoption. wiro is a small platform with a small community, its mcp server has near-zero installs, and there is no public pricing page to send a finance team to. if you hit an undocumented edge you will be filing a ticket rather than finding a stack overflow answer.

    pros
    • +dynamicprice and approximatelycost expose a job's cost before you run it
    • +billed only for completed runs — server errors and cancelled tasks are free
    • +media, llm, chat and realtime behind one api, unlike fal
    • +flat per-request pricing on its own endpoints — $0.006 for flux-2-dev at any resolution
    • +official mcp server, so an agent can drive the catalogue directly
    cons
    • 538 models against fal's 1,000+ and Replicate's community library
    • concurrency is capped at 10% of your account balance below $250
    • no public pricing page — wiro.ai/pricing returns a 404
    • small community, so undocumented edges mean a support ticket
  3. 3

    Replicate

    the community library, and the best story for your own model

    83/100

    verdictthe widest long tail and the only first-class path for packaging your own model, with private-deployment billing that charges for time your model spends doing nothing.

    best for
    teams running a fine-tuned or niche community model rather than a frontier one
    price
    $0.025 / image for flux.1 [dev]
    pricing note
    per-output on public models; private deployments are billed per second of hardware time, with idle and setup time billed separately except on fast-booting fine-tunes
    free tier
    no
    flux.1 [dev]
    $0.025 / image
    billing
    per output; per second for private deploys
    catalogue
    thousands, community-contributed
    llms too
    yes
    mcp server
    none published

    Replicate's public models are billed per output — flux.1 [dev] at $0.025 an image, directly comparable with fal — and behind that sits a community library nobody else matches. if the thing you need is an obscure upscaler, a specific fine-tune or a research checkpoint somebody wrapped last month, this is where it already is.

    Cog is the real differentiator. it is Replicate's open-source packaging tool, and it turns a model into an api server without you writing one. that makes Replicate the natural home for a team whose competitive advantage is a fine-tune rather than a prompt, and it is a materially better custom-model story than fal, Wiro or WaveSpeed offer.

    it also spans text as well as media, so the same account covers a chat model and an image model. that breadth is real, though for text specifically the dedicated inference providers price better.

    the trap is private deployments. those bill per second of hardware time and — the part that catches people — idle and setup time are billed separately, except on fast-booting fine-tunes. a private deployment left running between bursts costs money for doing nothing, which is exactly the failure mode per-output pricing exists to prevent. keep to public models and you avoid it entirely.

    pros
    • +the largest community model library in the category
    • +Cog makes deploying your own model a first-class path
    • +public models billed per output at $0.025/image, comparable with fal
    • +covers text models as well as media
    cons
    • private deployments bill idle and setup time separately
    • no exact catalogue count published, only 'thousands of models'
    • community models vary wildly in maintenance and documentation
    advertisement
  4. 4

    Runware

    the cheapest headline number in the category, and you cannot check it

    75/100

    verdictgenuinely aggressive pricing on a large catalogue, undercut by the fact that you cannot confirm what any specific model costs without signing up.

    best for
    high-volume image generation where you can benchmark your own model choice first
    price
    $0.0006 / image at the floor
    pricing note
    the $0.0006 to $0.24 range spans the whole image catalogue rather than any one model; the per-model price for flux.1 [dev] is not published on the pricing page, the model page or a dedicated url, and appears to require the in-app playground
    free tier
    yes
    flux.1 [dev]
    not publicly priced
    billing
    per output, varies by model and settings
    catalogue
    303 image, 38 llm listed
    llms too
    yes — 38 models
    mcp server
    yes — runware mcp

    Runware built its own inference stack — the Sonic Inference Engine — and prices off it rather than off commodity gpu rental, which is how the floor gets to $0.0006 an image. the pricing table lists 303 image models and 38 llm models, there is $2 of free credit on signup, and the company raised a $50m series a in early 2026, which suggests the number is a strategy rather than a stunt.

    there is a native mcp server referenced in the tools section of the pricing page, which puts Runware in the small group here that an agent can drive directly.

    and then the problem. we went looking for the price of one flux.1 [dev] image across the pricing page, a dedicated per-model url and the model explorer. the dedicated url 404s, the model page showed two unlabelled example figures without stating what resolution or step count produced them, and the explorer prints no prices at all. for a company whose entire market position is that it is the cheapest, not publishing a checkable per-model price is a strange decision.

    the practical consequence is that the $0.0006 floor is unfalsifiable from outside. it is almost certainly true of some model at some setting. whether it is true of the model you intend to run is something you find out after creating an account, which is the opposite of how the rest of this ranking behaves.

    pros
    • +lowest published price floor in the category at $0.0006/image
    • +own inference stack rather than resold gpu capacity
    • +$2 in free credits on signup
    • +native mcp server referenced on the pricing page
    cons
    • no public per-model price — flux.1 [dev] could not be priced from three separate pages
    • the quoted range spans the entire catalogue, so it prices nothing specific
    • confirming any real cost appears to require an account and the playground
  5. 5

    WaveSpeed AI

    half fal's price on the anchor, a fraction of the catalogue

    73/100

    verdictless than half what fal charges on the one model we can compare directly, on a catalogue small enough that you should check your models are on it first.

    best for
    a known, stable set of models where the per-image saving compounds
    price
    $0.012 / image for flux.1 [dev]
    pricing note
    per-output pay-as-you-go with no monthly commitment; cheapest listed model is Z Image Turbo at $0.005/image, and $1 of free credit is granted on signup with no card required
    free tier
    yes
    flux.1 [dev]
    $0.012 / image
    billing
    per output, pay-as-you-go
    catalogue
    40+ listed on pricing
    llms too
    yes
    mcp server
    none published

    on the anchor, WaveSpeed wins outright: $0.012 for a 1024x1024 flux.1 [dev] image against $0.025 a megapixel at fal and $0.025 an image at Replicate. pricing is per output with no monthly commitment, there is $1 of free credit without a card, and the platform covers text models as well as image and video.

    the vendor claims no cold starts and a median end-to-end generation around ten seconds. both are vendor claims we have not independently timed, and cold-start behaviour is exactly the kind of thing that degrades quietly under load, so treat them as a starting point.

    the catalogue is the constraint. the pricing page lists something over forty models and the vendor publishes no total count, which puts it an order of magnitude below fal on breadth. that is fine if the models you need are on it and a hard stop if they are not, so check before you plan a migration.

    custom model deployment exists only on the enterprise tier, and we could not confirm an openai-compatible endpoint or an mcp server from the vendor's own pages. one more thing worth naming: WaveSpeed publishes its own comparison posts ranking itself first against the other platforms in this list. nothing on this page comes from those.

    pros
    • +$0.012 per flux.1 [dev] image — less than half fal or Replicate
    • +covers image, video and text models
    • +$1 free credit on signup, no card required
    • +pay-as-you-go with no monthly commitment
    cons
    • catalogue is roughly forty listed models with no published total
    • custom deployment is enterprise-tier only
    • no cold-start or latency figure we could verify independently
  6. 6

    Novita AI

    one pricing page for text, image and video — minus the model everyone benchmarks

    69/100

    verdicta genuinely broad catalogue on a single pricing page, weakened by not carrying the base model this category is most often compared on.

    best for
    teams who want text, image and video on one bill and are not tied to a specific base model
    price
    $0.018 / image for flux.1 kontext dev in fast mode
    pricing note
    base flux.1 [dev] is not listed on the pricing page at all; the kontext variants are, at $0.0225 standard, $0.018 fast mode, $0.072 max and $0.36 pro. billing is per million tokens for text, per image for generation and per second for video
    free tier
    no
    flux.1 [dev]
    not carried
    billing
    per token, per image, per second
    catalogue
    200+ models
    llms too
    yes
    mcp server
    none published

    Novita puts more than two hundred models across text, image and video behind one account and one pricing page, with a sensible unit for each — per million tokens for llms, per image for generation, per second for video. for a team that wants one vendor relationship instead of three, that is the pitch, and it is a reasonable one.

    the anchor is where it comes apart. base flux.1 [dev] is not on the pricing page. what is listed is the Kontext family — $0.0225 for kontext dev, $0.018 in fast mode, $0.072 for max and $0.36 for pro — which are image editing models built on flux, not the base text-to-image checkpoint. that is not a pricing failure so much as a catalogue gap, but it does mean the most common benchmark in open image generation is not available here.

    the $0.018 fast-mode figure is competitive against the field, and if kontext-class editing is what you actually need this ranks better for you than its position suggests.

    we could not confirm a free tier, an openai-compatible endpoint, custom model deployment or an mcp server from Novita's own pages inside our fetch budget. the catalogue claim of two hundred plus models is the vendor's own, printed on the pricing page.

    pros
    • +text, image and video on one account and one pricing page
    • +200+ models claimed, with a sensible billing unit per modality
    • +$0.018 per image in fast mode is competitive on the kontext family
    cons
    • base flux.1 [dev] is not carried — only kontext variants
    • no free tier confirmable from the vendor's own pages
    • custom deployment and mcp support unconfirmed
  7. 7

    Segmind

    508 models, billed in gpu-seconds you cannot forecast

    64/100

    verdicta large catalogue and a real fine-tuning service, priced in a unit that makes the cost of a single image genuinely hard to predict.

    best for
    fine-tuning workloads where gpu-second billing is the natural unit anyway
    price
    $0.0072 / gpu-second for flux.1 [dev]
    pricing note
    serverless rate; dedicated cloud runs $0.0007 to $0.0031 per gpu-second. no per-image figure is printed for flux.1 [dev] — the model page shows a ~20.27s runtime alongside the rate but does not state a cost per image
    free tier
    no
    flux.1 [dev]
    $0.0072 / gpu-second
    billing
    per gpu-second + subscription credits
    catalogue
    508 models
    llms too
    yes
    mcp server
    none published

    Segmind lists 508 models spanning image, video and text, and it is one of the few platforms here with a built-in fine-tuning service rather than a bring-your-own-container story — the vendor suggests around twenty training images to start. for a team whose workflow is train, then serve, that consolidation has value.

    the pricing is the problem. flux.1 [dev] is billed at $0.0072 per gpu-second on serverless, and the model page shows a runtime figure of about 20.27 seconds without ever stating what an image costs. we are not going to multiply those two numbers and publish the result as a price, because a runtime that varies with step count and load is not a rate. the honest statement is the one Segmind makes: $0.0072 a gpu-second, and the cost of your image is however long the gpu takes.

    that unit is defensible for fine-tuning, where you are buying gpu time on purpose. it is a poor fit for per-image generation, and it is the single reason this sits last — every platform above it can answer the question 'what does this image cost' and this one structurally cannot.

    the structure around it is fragmented too. there is a pay-as-you-go tier referencing $10, subscription tiers running from $39 a month upward with monthly credit allowances, and separate per-gpu-second pricing for fine-tuning and dedicated endpoints, and reconciling which pool a given call draws from takes longer than it should. we could not confirm a free tier, an openai-compatible endpoint or an mcp server from the vendor's own pages.

    pros
    • +508 models across image, video and text
    • +built-in fine-tuning service, not just custom container hosting
    • +dedicated cloud rates drop to $0.0007 per gpu-second
    cons
    • no per-image price for flux.1 [dev] — cost depends on runtime
    • pricing split across pay-as-you-go, subscription credits and gpu-time tiers
    • the model page 404s at the obvious url, and the free tier is ambiguous

how this ranking was made

the anchor is one 1024x1024 image from flux.1 [dev], read from each platform's own model or pricing page on 21 august 2026. where a platform prints a different unit — per megapixel, per gpu-second, per request — the unit is reported as printed and not converted. multiplying a gpu-second rate by an advertised average runtime would produce a number that looks like a price and is not one.

where a figure could not be confirmed from the vendor's own page inside a fixed fetch budget, the entry says so rather than carrying a guess. that applies to Runware's flux.1 [dev] price, which we tried three separate pages for. an unpublished price is a finding about the platform, not a gap in the research.

Wiro's figures come from its own api and documentation rather than a marketing page, because it does not have a public pricing page — wiro.ai/pricing returns a 404 and prices live on individual model pages and in the api response.

no listicles were used for any figure here. platform comparison posts in this category are mostly written by the platforms themselves, and several of the vendors ranked below publish blog posts placing themselves first.

none of these platforms are operated by us, and none of the links are affiliate links.

our general methodology and disclosures →
was this useful?