verifier.org

best 16 ai video models

ranked on motion and prompt adherence, native audio, clip length, licence risk, and cost per second of finished video.

last reviewed 22 jul 2026 · 16 tools tested ·list curated by Onur Ozcanxin

the short version
best overallGemini Omni Flashalmost everything, until you hit the ten-second or 720p ceiling94/100runner-upSeedance 2.0multi-shot narrative work where director-level camera control matters90/100best free optionLTX-2.3self-hosted pipelines needing long clips, 4K, and audio in one pass73/100

video is where the money actually goes. at ten cents to seventy cents per second, a single minute of generated footage costs between six and forty dollars — so unlike image models, picking the wrong one here shows up on an invoice within a week.

native audio stopped being a feature and became table stakes. a year ago 'it generates its own soundtrack' was a headline; today most of this list produces synchronised dialogue, effects and ambience in a single pass, and artificial analysis maintains separate with-audio and without-audio leaderboards because the market has split. the holdouts — luma, the older open-weight models — now wear it as their defining weakness.

the other thing that changed is legal risk. seedance's launch drew cease-and-desists from disney and paramount skydance and a senate demand for shutdown; xai faces deepfake litigation and a multistate attorney-general letter. model choice is now partly a risk decision, which is the only reason adobe's mediocre model still has a market.

openai is not on this list as a recommendation. sora's consumer app closed in april and the api shuts down on 24 september 2026 — it is ranked here on merit with an explicit warning, not as something to build on.

advertisement
  1. 1

    Gemini Omni Flash

    first on every leaderboard, and one of the cheapest

    94/100

    verdictthe rare model that wins on quality and price at the same time — first on all four artificial analysis boards and first on lmarena, at a quarter of what google charges for veo 3.1.

    best for
    almost everything, until you hit the ten-second or 720p ceiling
    price
    $0.10 / second of 720p, direct from google
    pricing note
    fal charges roughly $0.13/second for the same thing — going direct saves about 23%
    free tier
    yes
    access
    closed api (gemini, vertex) + apps
    native audio
    yes — single pass
    max duration
    10 seconds
    max resolution
    720p
    license
    commercial via api terms

    this is the clearest result in either of our model categories. gemini omni flash is first on artificial analysis for text-to-video and image-to-video, with and without audio, and first on lmarena across half a million votes. two independent vote pools agreeing this completely is unusual.

    and it costs ten cents a second. google's own veo 3.1 standard is forty cents a second and now ranks tenth on the same board. google has undercut itself by 4× and beaten itself on quality at the same time, which tells you something about how fast this category is moving.

    audio is generated natively in the same pass — dialogue, effects and ambience, synchronised. it could not top the with-audio leaderboards otherwise.

    the limits are real and worth planning around: ten seconds maximum, 720p maximum, and it is still a preview. google also documents that character consistency degrades across cuts and pans, and that video references up to three seconds are accepted by the api but not correctly processed yet. for short-form social, ads and b-roll that is fine. for anything longer you are stitching.

    pros
    • +#1 on all four artificial analysis boards and on lmarena
    • +$0.10/second undercuts almost everything above 720p
    • +native synchronised audio in a single pass
    • +available through google ai studio, gemini api, flow, fal and runway
    cons
    • ten-second maximum per generation
    • 720p ceiling — no 1080p or 4K
    • character consistency slips across cuts and pans
  2. 2

    Seedance 2.0

    the best cinematic control, wrapped in a copyright fight

    90/100

    verdictsecond on both leaderboards and the best model here for actual filmmaking — but it is three times the price of the leader and carries genuine legal baggage.

    best for
    multi-shot narrative work where director-level camera control matters
    price
    $0.30 / second at 720p with audio
    pricing note
    $0.68/second at 1080p; the fast variant is $0.24; audio costs nothing extra because you are charged for it either way
    free tier
    no
    access
    closed api (volcano, byteplus, fal)
    native audio
    yes — included in the price
    max duration
    15 seconds
    max resolution
    1080p
    license
    commercial via api terms

    seedance 2.0 is the strongest model on this list for multi-shot storytelling: physics that hold up, director-level camera control, and shots that feel composed rather than generated. it sits second on both artificial analysis and lmarena, behind only gemini omni flash.

    it does fifteen seconds against omni flash's ten, and up to 1080p against omni flash's 720p, so for anything that needs length or resolution it is the obvious step up.

    the price is the problem. $0.30 a second at 720p is triple the leader; $0.68 at 1080p is the most expensive tier of any mainstream model. and a detail worth knowing: you are charged the same whether audio generation is on or off, so switching it off saves you nothing at all.

    then there is the legal position. seedance's launch drew a cease-and-desist from disney, an infringement allegation from paramount skydance covering star trek, south park and dora, a statement from the mpa's chief executive accusing bytedance of unauthorised use of us copyrighted works on a massive scale, and a senate letter demanding shutdown. bytedance paused the international rollout. none of that makes the model worse, but it is a real consideration for anything client-facing.

    pros
    • +#2 on both artificial analysis and lmarena
    • +best multi-shot and camera control in the category
    • +fifteen seconds at up to 1080p, with native audio
    • +reference-to-video accepts nine images, at a 0.6× discount
    cons
    • 3× the price of the leader; $0.68/second at 1080p
    • disabling audio saves you nothing — same price either way
    • active copyright disputes with major studios
  3. 3

    Wan 2.7

    near-frontier quality at a third of the price

    87/100

    verdictthird on the with-audio leaderboard at a third of seedance's price, with four generation modes including instruction-based editing — the best value in video, full stop.

    best for
    high-volume production where cost per finished second is the constraint
    price
    $0.10 / second, all resolutions
    pricing note
    flat rate across 720p and 1080p — a ten-second 1080p clip costs $1.00
    free tier
    no
    access
    closed api (fal, alibaba cloud)
    native audio
    yes — preserve or regenerate
    max duration
    15 seconds
    max resolution
    1080p
    license
    commercial via api terms

    alibaba's tongyi lab shipped wan 2.7 in april with four modes: text-to-video, image-to-video, reference-to-video, and instruction-based video editing. that last one is rarer than it should be and genuinely useful.

    it ranks third on artificial analysis for text-to-video with audio — above kling, above veo 3.1 — at a flat ten cents a second regardless of resolution. seedance sits two places higher and costs three to seven times more. if you are producing volume, this is where the maths lands.

    it does 2–15 seconds, 720p and 1080p, all the standard aspect ratios, nine-grid multi-image input, first-and-last-frame control and character reference. native audio can either preserve the original or regenerate it.

    one correction worth making loudly, because a lot of coverage gets it wrong: wan 2.7 is not open weights. alibaba's public releases stop at wan 2.2, which is apache 2.0. the sites describing 2.7 as a downloadable apache model are wrong, and if you need weights you are getting 2.2, not 2.7.

    pros
    • +#3 on artificial analysis text-to-video with audio
    • +flat $0.10/second regardless of resolution
    • +four modes including instruction-based video editing
    • +first-and-last-frame control and character reference
    cons
    • not open weights despite widespread claims — 2.2 is the open line
    • documentation is thin compared to western vendors
    • fifteen-second ceiling
    advertisement
  4. 4

    Happy Horse 1.1

    the most underrated model in the category

    85/100

    verdictfourth on the with-audio leaderboard at eighteen cents a second for 1080p, with multilingual lip sync — it beats models costing more than twice as much and almost nobody has heard of it.

    best for
    multilingual dialogue video where lip sync has to actually land
    price
    $0.14 / second at 720p, $0.18 at 1080p
    pricing note
    from alibaba's taotian group — a different team from the tongyi lab that makes wan
    free tier
    no
    access
    closed api (fal, alibaba cloud)
    native audio
    yes — with multilingual lip sync
    max duration
    15 seconds
    max resolution
    1080p
    license
    commercial via api terms

    happy horse comes from alibaba's taotian group rather than the tongyi lab behind wan, which is why it shows up on leaderboards as a separate vendor. it launched in june and immediately placed fourth on artificial analysis for text-to-video with audio.

    for context on what that means: kling's 4K tier costs $0.42 a second and ranks below it. happy horse does 1080p with native audio for $0.18.

    the standout feature is multilingual lip sync — mouth movement matched to speech across languages. for dubbed content, localised ads, or any talking-head format that has to ship in more than one market, that is the specific thing most models do badly.

    the catch is entirely commercial rather than technical: minimal brand recognition, thin western distribution outside aggregators, and no 4K. if you can get past not recognising the name, the price-performance here is close to unbeatable.

    pros
    • +#4 on artificial analysis text-to-video with audio
    • +multilingual lip sync that actually tracks speech
    • +$0.18/second for 1080p with audio
    • +fifteen seconds at 24fps
    cons
    • no 4K option
    • almost no brand recognition or community
    • western availability limited to aggregators
  5. 5

    Kling 3.0

    the only native 4K video, and the deepest shot-level control

    82/100

    verdictthe best control surface in video and the only genuinely native 4K — priced well above its leaderboard position, which is the trade you're making.

    best for
    storyboarded sequences where each shot needs its own duration, framing and camera move
    price
    $0.084 / second standard with audio off; $0.42 for native 4K
    pricing note
    v3 pro image-to-video is $0.112 audio off, $0.168 audio on, $0.196 with voice control
    free tier
    yes
    access
    closed api (kling, fal) + web app
    native audio
    yes — five languages
    max duration
    15 seconds
    max resolution
    4K native
    license
    commercial via api terms

    kuaishou's kling 3.0 landed in february and its differentiator is control. you can storyboard a sequence and specify per shot: duration, shot size, perspective, camera movement. element consistency via reference uploads holds subjects across those shots. nothing else here gives you that much direction.

    it is also the only model on this list producing native 4K — not an upscale — at $0.42 a second.

    native audio covers english, chinese, japanese, korean and spanish including regional accents, though at the 4K tier audio narrows to chinese and english with everything else auto-translated to english. that is an odd limitation to hit late in a project.

    the honest problem is value. it sits sixth and ninth on the artificial analysis text-to-video board while costing more than models above it. you are paying for the control surface and the 4K, not for output quality. one more thing: the widely-repeated claim that kling 3.0 runs at 60fps does not appear in kuaishou's own announcement — the 4K endpoint is confirmed, the frame rate is not.

    pros
    • +only true native 4K video generation here
    • +per-shot control over duration, framing and camera movement
    • +five-language native audio with regional accents
    • +element consistency across a storyboarded sequence
    cons
    • ranks below cheaper models on quality leaderboards
    • 4K audio drops to chinese and english only
    • the '60fps' claim is unverified — don't plan around it
  6. 6

    Veo 3.1

    best audio engineering, overtaken by its own stablemate

    80/100

    verdictstill the best-sounding model here and the only route to two and a half minutes of continuous output — but google's own omni flash beats it on quality at a quarter of the price.

    best for
    long-form assembly and enterprise work that needs indemnity and 48khz audio
    price
    $0.40 / second standard; Fast $0.10 at 720p; Lite $0.05
    pricing note
    4K is $0.60/second on standard; there is no veo 4 — anything you have read about one is speculation
    free tier
    yes
    access
    closed api (gemini, vertex) + flow
    native audio
    yes — 48khz
    max duration
    8s native, ~148s via extension
    max resolution
    4K
    license
    commercial, enterprise indemnity

    veo 3.1's audio is genuinely the best engineered on this list: 48khz, with dialogue lip-synced properly, effects and ambience. if the output is going anywhere near a professional audio chain, that number matters.

    video extension is its other real advantage. native generations are four, six or eight seconds, but chaining up to twenty extensions gets you to roughly 148 seconds of continuous video — far beyond anything else here, all of which cap out between ten and twenty seconds.

    it does 4K, comes with google's enterprise trust and indemnity story, and integrates with flow.

    the awkward part is that gemini omni flash — also google's — ranks first on the board where veo 3.1 now ranks tenth, at $0.10 a second against $0.40. if you are on veo 3.1 standard today and don't need the length or the 4K, you are paying four times over for a lower-ranked model. the fast tier at $0.10 is the more defensible choice.

    pros
    • +48khz audio — the best sound quality in the category
    • +extension chains to roughly 148 seconds of continuous video
    • +native 4K at $0.60/second
    • +enterprise indemnity and google support
    cons
    • $0.40/second standard is 4× google's own better-ranked model
    • only 4–8 seconds per native generation
    • now ranks tenth on artificial analysis text-to-video with audio
  7. 7

    Grok Imagine Video 1.5

    excellent price-to-quality, serious brand-safety baggage

    76/100

    verdictthird on both image-to-video boards at eight cents a second — a genuinely strong model attached to a genuinely serious reputational problem.

    best for
    cost-sensitive image-to-video where 720p is enough
    price
    $0.08 / second direct from xai
    pricing note
    fal charges $0.08 at 480p and $0.14 at 720p — about 75% more than xai direct at 720p
    free tier
    no
    access
    closed api (xai, fal)
    native audio
    yes — with lip sync
    max duration
    15 seconds
    max resolution
    720p
    license
    commercial via api terms

    on the numbers this is one of the best-value models in the category: third on artificial analysis for image-to-video both with and without audio, beating models that cost three to four times more, with native synchronised audio including lip-synced dialogue.

    buy it direct from xai rather than through fal — xai charges $0.08 a second flat, while fal charges $0.14 at 720p. that is a 75% markup for the same model.

    the ceiling is 720p, with no 1080p option at all, and up to fifteen seconds.

    the part that belongs in any honest review: xai faces multiple lawsuits over non-consensual sexualised deepfakes, reporting that the video tool produced sexualised imagery of a named public figure without being explicitly asked, a multistate attorney-general letter, and an amsterdam court threatening €100,000-a-day fines. safeguards added in january were reported to apply only to non-paying users. whatever you conclude about the model, putting it in a consumer-facing product is a decision to take with your legal team rather than your engineering one.

    pros
    • +#3 on both artificial analysis image-to-video boards
    • +$0.08/second direct is excellent value
    • +native synchronised audio with lip sync
    • +fifteen-second clips
    cons
    • 720p ceiling, no 1080p
    • active deepfake litigation and regulatory pressure
    • fal charges ~75% over xai direct at 720p
  8. 8

    PixVerse V6

    the value pick for silent video

    75/100

    verdictstatistically tied for fourth on image-to-video without audio, at nine cents a second for 1080p — if you don't need generated sound, this is the efficient frontier.

    best for
    high-volume 1080p b-roll where you're adding your own soundtrack anyway
    price
    $0.09 / second at 1080p without audio
    pricing note
    $0.115 with audio; cheaper tiers run from $0.025/second at 360p — unlike seedance, turning audio off genuinely saves money
    free tier
    yes
    access
    closed api (fal) + web app
    native audio
    yes — but the weak part
    max duration
    15 seconds
    max resolution
    1080p
    license
    commercial via api terms

    pixverse v6 ranks fourth on artificial analysis for image-to-video without audio, effectively tied with grok 1.5, at $0.09 a second for 1080p. for silent footage that is the best quality-per-dollar on this list.

    the pricing model is also more honest than most: audio is a per-second uplift you can decline. seedance charges you for audio whether you use it or not; here, switching it off actually reduces the bill.

    it ships more than twenty cinema camera controls and a multi-shot engine, and handles one to fifteen seconds at up to 1080p.

    the tell is in the leaderboard split: its with-audio ranking is markedly weaker than its without-audio ranking, which points at the audio component being the weak part. treat it as a strong silent-video model that happens to offer sound, rather than an audio-video model.

    pros
    • +#4 on image-to-video without audio at $0.09/second for 1080p
    • +audio is optional and genuinely cheaper when off
    • +20+ cinema camera controls and a multi-shot engine
    • +tiers from $0.025/second for drafts
    cons
    • audio quality is clearly the weak component
    • frame rate not published
    • no 4K
  9. 9

    LTX-2.3

    the best open-weight video model, with a licence that isn't what you think

    73/100

    verdicttwenty-second clips at up to 4K with genuine single-pass audio, running on your own hardware — just don't believe anyone who tells you it's apache 2.0.

    best for
    self-hosted pipelines needing long clips, 4K, and audio in one pass
    price
    free to self-host; $0.06 / second at 1080p on fal
    pricing note
    4K is $0.24/second hosted; the licence requires a paid commercial agreement above $10M annual revenue
    free tier
    yes
    access
    open weights + hosted api
    native audio
    yes — single pass
    max duration
    20 seconds
    max resolution
    4K
    license
    community licence, $10M revenue gate

    ltx-2.3 is a 22-billion-parameter audio-video model from lightricks, and it is comfortably the strongest open-weight option here: leader among open weights on artificial analysis' image-to-video with audio board, twenty-second clips — the longest of any mainstream model — up to 4K, and 24 or 48fps.

    it generates synchronised video and audio in a single forward pass, which most open models still can't do at all.

    hosted on fal it is also cheap: $0.06 a second at 1080p, $0.24 at 4K, with a fast tier at two-thirds of that.

    now the licence, because this is the single most misreported fact in the category — including on fal's own model pages, which state the model is apache 2.0. it is not. the actual licence file is the ltx-2 community license agreement, and it requires organisations with at least $10 million in annual revenue to obtain a paid commercial licence. it also prohibits use in products competing with lightricks' own, and specifies liquidated damages of double the licence fee for unauthorised commercial use. read the licence file, not the blog post.

    pros
    • +best open-weight video model by a clear margin
    • +twenty-second clips — the longest here
    • +4K output with genuine single-pass audio
    • +cheap hosted option at $0.06/second for 1080p
    cons
    • not apache 2.0 — $10M revenue gate and anti-competition clause
    • quality still well behind the closed frontier
    • self-hosting a 22b video model needs serious gpu capacity
  10. 10

    Runway Gen-4.5

    the toolchain is the product now

    70/100

    verdictthe base model no longer competes, but aleph video-to-video and act-two performance capture still have no real equivalent — you're buying the workflow, not the weights.

    best for
    editing workflows, performance capture, and video-to-video rather than raw generation
    price
    $0.12 / second (12 credits at $0.01)
    pricing note
    gen-3 alpha turbo and gen-4 aleph are being sunset on 30 july 2026 — check what your pipeline depends on
    free tier
    yes
    access
    closed api + web app
    native audio
    yes — on gen-4.5
    max duration
    not published
    max resolution
    720p native, 4K upscale
    license
    commercial on paid plans

    runway's gen-4.5 generates natively at 720p, which in a field that reached 4K is a real limitation, and it appears in neither leaderboard's top ten. as a text-to-video model it has been passed.

    what runway still has is everything around the model: aleph for video-to-video transformation, act-two for performance capture, a genuine editing timeline, and 4K upscaling. those are workflow tools that the model labs mostly haven't built, and for teams doing post rather than generation they remain the reason to be here.

    runway has also become substantially a reseller. its pricing page now sells seedance 2.0, veo 3, veo 3.1, gemini omni flash and happy horse alongside its own model. that is a sensible read of where things are going, but it does mean the differentiator is the interface rather than the output.

    one urgent operational note: gen-3 alpha turbo and gen-4 aleph are being sunset on 30 july 2026. if anything you run depends on those endpoints, that is days away, not months.

    pros
    • +aleph video-to-video has no real equivalent elsewhere
    • +act-two performance capture and a real editing timeline
    • +aggregates seedance, veo, omni flash and happy horse in one place
    • +native audio on gen-4.5, generated in the same pass
    cons
    • native generation is only 720p
    • absent from both leaderboards' top ten
    • gen-3 alpha turbo and gen-4 aleph sunset 30 july 2026
  11. 11

    Luma Ray3.2

    the only one that speaks a professional colour pipeline

    66/100

    verdictnative 16-bit HDR and EXR export in ACES is unique here and genuinely matters to post houses — but it generates no audio in a year when everything else does.

    best for
    vfx and finishing work that has to survive a colour grade
    price
    ~$1.20 per 5-second 1080p SDR generation
    pricing note
    credit-based and luma doesn't publish a per-credit rate; HDR is ~2× and HDR+EXR ~3×
    free tier
    yes
    access
    closed api + web app
    native audio
    no — third-party integration
    max duration
    20 seconds
    max resolution
    1080p, 16-bit HDR
    license
    commercial on paid plans

    first, a correction: dream machine, ray2, ray3 and ray3.14 are all superseded. ray3.2 arrived in june and anything recommending luma dream machine is out of date.

    luma's real differentiator is the colour pipeline. native 16-bit HDR and EXR export in ACES2065-1 is something no other model on this list offers, and for a vfx shop that has to composite and grade generated footage alongside real plates, that is the difference between usable and not.

    it also offers up to sixteen keyframes per clip, which is the finest-grained motion direction available anywhere here, and modify video v2 handles up to twenty seconds at 1080p.

    the gap is audio. luma generates none, integrating third-party tools instead, in a year when single-pass native audio became standard. pricing is also opaque — luma doesn't publish a per-credit dollar value, so the roughly $1.20 per five-second generation figure is an approximation rather than a quote.

    pros
    • +native 16-bit HDR and EXR/ACES export — unique here
    • +up to sixteen keyframes for fine motion control
    • +twenty seconds at 1080p on modify video
    • +the right choice for professional finishing pipelines
    cons
    • no native audio at all
    • no published per-credit pricing
    • absent from leaderboard top tens
  12. 12

    MiniMax Hailuo 2.3

    good motion, impossible pricing

    62/100

    verdicta genuinely capable model behind the least transparent pricing in the category — $1,000 minimum packages that expire monthly are hard to recommend to anyone small.

    best for
    teams already committed to minimax who can absorb the package minimums
    price
    opaque — sold in 'video points', from $1,000 per package
    pricing note
    minimax publishes no per-second rate; every per-second figure in circulation comes from resellers, and points expire after a month
    free tier
    no
    access
    closed api + consumer apps
    native audio
    unconfirmed in official docs
    max duration
    not published
    max resolution
    1080p
    license
    commercial via api terms

    hailuo has a strong reputation for motion quality and a large consumer following, and the model itself is competitive.

    the pricing is the problem and it is worth being blunt about it. minimax sells 'video points' in packages starting at $1,000 for 3,760 points, rising to $6,000 — and the points expire after a month. the documentation gives deduction examples rather than rates, so working out what a project costs means reverse-engineering it from worked examples.

    every per-second figure you will find for hailuo comes from third-party resellers, not from minimax. we are not going to publish those as if they were the vendor's prices.

    there are consumer plans from $9.99 and a developer token plan from $10 a month, which are more approachable, but the api story for anyone building at scale involves those package minimums. capable model, hostile commercial model.

    pros
    • +strong motion quality and physics
    • +large consumer following and community knowledge
    • +consumer plans start at $9.99/month
    • +developer token plan from $10/month
    cons
    • no published per-second pricing at all
    • api packages start at $1,000 and expire after a month
    • native audio status not confirmed in official docs
  13. 13

    Adobe Firefly Video

    the risk-averse choice, and it knows it

    59/100

    verdictabsent from every leaderboard, and still the obvious pick if your legal team has watched what happened to bytedance this year.

    best for
    enterprises that need licensed training data and indemnification more than they need quality
    price
    $9.99 / month consumer; api from a ~$1,000 / month enterprise floor
    pricing note
    consumer plans do not include api access; firefly services is a separate enterprise contract
    free tier
    yes
    access
    subscription; api at enterprise tier
    native audio
    no — separate paid features
    max duration
    not published
    max resolution
    1080p
    license
    licensed data, indemnified

    firefly video is not a competitive model on output quality and does not appear in any leaderboard we checked.

    it is the only major video model with a clean commercial-safety story: trained on licensed and adobe stock content, with indemnification available. in a year when seedance drew cease-and-desists from disney and paramount and xai collected deepfake lawsuits, that is a product feature with real commercial value.

    be precise about the audio claim, though. adobe offers audio translation, lip sync and sound-effect generation as separate credit-consuming features. that is not the same as single-pass native audio generation, and vendors are increasingly blurring the two.

    the commercial structure is awkward for anyone small: consumer plans from $9.99 to $199.99 a month don't include api access at all, and the firefly services api starts around a thousand dollars a month. like the image model, firefly video increasingly routes to partner models rather than competing directly.

    pros
    • +licensed training data with enterprise indemnification
    • +deep creative cloud and premiere integration
    • +the defensible choice when legal risk is the priority
    • +consumer plans start at $9.99/month
    cons
    • absent from every quality leaderboard
    • api access starts around $1,000/month
    • audio features are bolted on, not single-pass native
  14. 14

    HunyuanVideo 1.5

    runs on one consumer gpu, banned in three jurisdictions

    55/100

    verdict8.3 billion parameters on a single rtx 4090 is a real achievement, and the licence excludes three major markets outright — check your geography before your gpu.

    best for
    fine-tuning research outside the eu, uk and south korea
    price
    free to self-host
    pricing note
    the licence expressly does not apply in the EU, the UK or South Korea — read it before you deploy
    free tier
    yes
    access
    open weights, self-hosted
    native audio
    no — separate foley model
    max duration
    ~5 seconds
    max resolution
    720p
    license
    tencent community — excludes EU/UK/KR

    hunyuanvideo 1.5 fits in 8.3 billion parameters and generates a clip in about seventy-five seconds on a single rtx 4090. for a research team or a solo developer without a cluster, that accessibility is the point.

    it is still widely used as a fine-tuning base, and the weights are genuinely downloadable.

    but the base model is silent — audio comes from a separate hunyuanvideo-foley model — and it tops out at roughly five-second clips at 720p. against ltx-2.3's twenty seconds at 4K with single-pass audio, it is a generation behind for open-weight work, and tencent has shipped nothing new on this line in 2026.

    the licence deserves a careful read and is the reason it sits this low. it is the tencent hunyuan community licence, not an osi-approved open-source licence: organisations above 100 million monthly active users must request separate permission at tencent's discretion, outputs may not be used to train competing models, attribution is required, and — the clause most people miss — the licence expressly does not apply in the european union, the united kingdom or south korea. for a european company that is a hard blocker, not a footnote.

    pros
    • +runs on a single consumer gpu
    • +genuinely downloadable weights
    • +still a solid fine-tuning base
    • +free to self-host
    cons
    • licence does not apply in the eu, uk or south korea
    • silent — audio needs a separate model
    • ~5 seconds at 720p, a generation behind ltx-2.3
  15. 15

    OpenAI Sora 2

    still good, shutting down on 24 september 2026

    40/100

    verdictsora 2 pro still ranks fifth on lmarena, and the api shuts down on 24 september 2026 — do not start anything here.

    best for
    nothing new — this is a migration notice, not a recommendation
    price
    ~$0.10 / second standard, ~$0.30 pro
    pricing note
    openai's pricing page blocks automated access; these figures are secondary-sourced and largely academic now
    free tier
    no
    access
    closed api — retiring sept 2026
    native audio
    yes
    max duration
    not published
    max resolution
    1024p on pro
    license
    commercial until shutdown

    this entry exists as a warning rather than a recommendation. openai is retiring sora in two stages: the consumer web and app experiences closed on 26 april 2026, and the api shuts down on 24 september 2026. developers on the videos api and sora 2 model aliases were formally notified in march.

    the model itself is not the problem. sora 2 pro still sits fifth on lmarena, ahead of every veo variant, and its native audio was genuinely category-defining when it launched.

    but a model with a published end-of-life two months out cannot be recommended for any new integration, and its low placement here reflects exactly that. capability is not the same as viability.

    if you are already on it, treat migration as urgent. gemini omni flash is cheaper and ranks higher; seedance 2.0 and wan 2.7 are the closest matches for longer or higher-resolution work.

    pros
    • +still #5 on lmarena text-to-video
    • +native audio was category-defining at launch
    • +strong physics and prompt adherence
    cons
    • api shuts down 24 september 2026
    • consumer apps already closed in april
    • no successor announced — openai has redirected the effort to research
  16. 16

    Pika 2.5

    fun effects, no longer a serious production tool

    38/100

    verdictstill enjoyable for social-first content, but it appears in no leaderboard and has no real api — it has fallen out of the competitive tier.

    best for
    playful social formats and effects, not production or api work
    price
    free tier; paid from $8 / month
    pricing note
    commercial rights require a paid plan; there is no meaningful public developer api
    free tier
    yes
    access
    subscription consumer app
    native audio
    yes
    max duration
    not published
    max resolution
    1080p
    license
    commercial on paid plans

    pika 2.5 handles physics interactions and synchronised audio, and its effects-driven approach is genuinely fun for social formats. that was enough to matter in 2024.

    it is now absent from every leaderboard we consulted, and there is no meaningful public developer api — it is a subscription consumer app.

    the free tier gives eighty credits a month and paid plans start at $8 with commercial rights, which makes it cheap to play with.

    we would not build anything on it. a widely-referenced 'pika 3.0' appears to be projected rather than shipped, so we're not ranking on it.

    pros
    • +cheap, with a usable free tier
    • +fun effects-driven approach for social content
    • +commercial rights from $8/month
    • +synchronised audio and physics interactions
    cons
    • absent from every leaderboard
    • no meaningful public developer api
    • 'pika 3.0' appears projected, not shipped

how this ranking was made

hands-on prompting around the things that actually break: a physics shot, a dialogue shot to check lip sync, a camera move, and a pair where the same character has to survive a cut. qualitative, and used to sanity-check the leaderboards rather than to produce a score of our own.

cost per second is the vendor's published rate at the resolution named in the entry, taken from the vendor's pricing page or the fal model page on the review date. video pricing is tiered by resolution and often by whether audio is on — we state which tier the number refers to, because a single blended 'price per second' for this category would be fiction.

native audio means generated in the same pass as the video. a model that hands off to a separate sound-effects or lip-sync product is marked as not having it, however good the combined result is.

we read the actual licence file. three of the most widely-cited 'apache 2.0' claims in video are wrong — ltx-2 ships a community licence with a $10m revenue gate and an anti-competition clause, wan's public weights stop at 2.2, and tencent's hunyuan licence explicitly does not apply in the eu, the uk or south korea.

we cross-check against the artificial analysis and lmarena video leaderboards. both currently agree on the top two, which is a stronger signal than either alone.

our general methodology and disclosures →
was this useful?