video is where the money actually goes. at ten cents to seventy cents per second, a single minute of generated footage costs between six and forty dollars — so unlike image models, picking the wrong one here shows up on an invoice within a week.
native audio stopped being a feature and became table stakes. a year ago 'it generates its own soundtrack' was a headline; today most of this list produces synchronised dialogue, effects and ambience in a single pass, and artificial analysis maintains separate with-audio and without-audio leaderboards because the market has split. the holdouts — luma, the older open-weight models — now wear it as their defining weakness.
the other thing that changed is legal risk. seedance's launch drew cease-and-desists from disney and paramount skydance and a senate demand for shutdown; xai faces deepfake litigation and a multistate attorney-general letter. model choice is now partly a risk decision, which is the only reason adobe's mediocre model still has a market.
openai is not on this list as a recommendation. sora's consumer app closed in april and the api shuts down on 24 september 2026 — it is ranked here on merit with an explicit warning, not as something to build on.
advertisement
1
Gemini Omni Flash
first on every leaderboard, and one of the cheapest
94/100
verdictthe rare model that wins on quality and price at the same time — first on all four artificial analysis boards and first on lmarena, at a quarter of what google charges for veo 3.1.
best for
almost everything, until you hit the ten-second or 720p ceiling
price
$0.10 / second of 720p, direct from google
pricing note
fal charges roughly $0.13/second for the same thing — going direct saves about 23%
free tier
yes
access
closed api (gemini, vertex) + apps
native audio
yes — single pass
max duration
10 seconds
max resolution
720p
license
commercial via api terms
this is the clearest result in either of our model categories. gemini omni flash is first on artificial analysis for text-to-video and image-to-video, with and without audio, and first on lmarena across half a million votes. two independent vote pools agreeing this completely is unusual.
and it costs ten cents a second. google's own veo 3.1 standard is forty cents a second and now ranks tenth on the same board. google has undercut itself by 4× and beaten itself on quality at the same time, which tells you something about how fast this category is moving.
audio is generated natively in the same pass — dialogue, effects and ambience, synchronised. it could not top the with-audio leaderboards otherwise.
the limits are real and worth planning around: ten seconds maximum, 720p maximum, and it is still a preview. google also documents that character consistency degrades across cuts and pans, and that video references up to three seconds are accepted by the api but not correctly processed yet. for short-form social, ads and b-roll that is fine. for anything longer you are stitching.
pros
+#1 on all four artificial analysis boards and on lmarena
+$0.10/second undercuts almost everything above 720p
+native synchronised audio in a single pass
+available through google ai studio, gemini api, flow, fal and runway
the best cinematic control, wrapped in a copyright fight
90/100
verdictsecond on both leaderboards and the best model here for actual filmmaking — but it is three times the price of the leader and carries genuine legal baggage.
best for
multi-shot narrative work where director-level camera control matters
price
$0.30 / second at 720p with audio
pricing note
$0.68/second at 1080p; the fast variant is $0.24; audio costs nothing extra because you are charged for it either way
free tier
no
access
closed api (volcano, byteplus, fal)
native audio
yes — included in the price
max duration
15 seconds
max resolution
1080p
license
commercial via api terms
seedance 2.0 is the strongest model on this list for multi-shot storytelling: physics that hold up, director-level camera control, and shots that feel composed rather than generated. it sits second on both artificial analysis and lmarena, behind only gemini omni flash.
it does fifteen seconds against omni flash's ten, and up to 1080p against omni flash's 720p, so for anything that needs length or resolution it is the obvious step up.
the price is the problem. $0.30 a second at 720p is triple the leader; $0.68 at 1080p is the most expensive tier of any mainstream model. and a detail worth knowing: you are charged the same whether audio generation is on or off, so switching it off saves you nothing at all.
then there is the legal position. seedance's launch drew a cease-and-desist from disney, an infringement allegation from paramount skydance covering star trek, south park and dora, a statement from the mpa's chief executive accusing bytedance of unauthorised use of us copyrighted works on a massive scale, and a senate letter demanding shutdown. bytedance paused the international rollout. none of that makes the model worse, but it is a real consideration for anything client-facing.
pros
+#2 on both artificial analysis and lmarena
+best multi-shot and camera control in the category
+fifteen seconds at up to 1080p, with native audio
+reference-to-video accepts nine images, at a 0.6× discount
cons
−3× the price of the leader; $0.68/second at 1080p
−disabling audio saves you nothing — same price either way
verdictthird on the with-audio leaderboard at a third of seedance's price, with four generation modes including instruction-based editing — the best value in video, full stop.
best for
high-volume production where cost per finished second is the constraint
price
$0.10 / second, all resolutions
pricing note
flat rate across 720p and 1080p — a ten-second 1080p clip costs $1.00
free tier
no
access
closed api (fal, alibaba cloud)
native audio
yes — preserve or regenerate
max duration
15 seconds
max resolution
1080p
license
commercial via api terms
alibaba's tongyi lab shipped wan 2.7 in april with four modes: text-to-video, image-to-video, reference-to-video, and instruction-based video editing. that last one is rarer than it should be and genuinely useful.
it ranks third on artificial analysis for text-to-video with audio — above kling, above veo 3.1 — at a flat ten cents a second regardless of resolution. seedance sits two places higher and costs three to seven times more. if you are producing volume, this is where the maths lands.
it does 2–15 seconds, 720p and 1080p, all the standard aspect ratios, nine-grid multi-image input, first-and-last-frame control and character reference. native audio can either preserve the original or regenerate it.
one correction worth making loudly, because a lot of coverage gets it wrong: wan 2.7 is not open weights. alibaba's public releases stop at wan 2.2, which is apache 2.0. the sites describing 2.7 as a downloadable apache model are wrong, and if you need weights you are getting 2.2, not 2.7.
pros
+#3 on artificial analysis text-to-video with audio
+flat $0.10/second regardless of resolution
+four modes including instruction-based video editing
+first-and-last-frame control and character reference
cons
−not open weights despite widespread claims — 2.2 is the open line
−documentation is thin compared to western vendors
verdictfourth on the with-audio leaderboard at eighteen cents a second for 1080p, with multilingual lip sync — it beats models costing more than twice as much and almost nobody has heard of it.
best for
multilingual dialogue video where lip sync has to actually land
price
$0.14 / second at 720p, $0.18 at 1080p
pricing note
from alibaba's taotian group — a different team from the tongyi lab that makes wan
free tier
no
access
closed api (fal, alibaba cloud)
native audio
yes — with multilingual lip sync
max duration
15 seconds
max resolution
1080p
license
commercial via api terms
happy horse comes from alibaba's taotian group rather than the tongyi lab behind wan, which is why it shows up on leaderboards as a separate vendor. it launched in june and immediately placed fourth on artificial analysis for text-to-video with audio.
for context on what that means: kling's 4K tier costs $0.42 a second and ranks below it. happy horse does 1080p with native audio for $0.18.
the standout feature is multilingual lip sync — mouth movement matched to speech across languages. for dubbed content, localised ads, or any talking-head format that has to ship in more than one market, that is the specific thing most models do badly.
the catch is entirely commercial rather than technical: minimal brand recognition, thin western distribution outside aggregators, and no 4K. if you can get past not recognising the name, the price-performance here is close to unbeatable.
pros
+#4 on artificial analysis text-to-video with audio
+multilingual lip sync that actually tracks speech
the only native 4K video, and the deepest shot-level control
82/100
verdictthe best control surface in video and the only genuinely native 4K — priced well above its leaderboard position, which is the trade you're making.
best for
storyboarded sequences where each shot needs its own duration, framing and camera move
price
$0.084 / second standard with audio off; $0.42 for native 4K
pricing note
v3 pro image-to-video is $0.112 audio off, $0.168 audio on, $0.196 with voice control
free tier
yes
access
closed api (kling, fal) + web app
native audio
yes — five languages
max duration
15 seconds
max resolution
4K native
license
commercial via api terms
kuaishou's kling 3.0 landed in february and its differentiator is control. you can storyboard a sequence and specify per shot: duration, shot size, perspective, camera movement. element consistency via reference uploads holds subjects across those shots. nothing else here gives you that much direction.
it is also the only model on this list producing native 4K — not an upscale — at $0.42 a second.
native audio covers english, chinese, japanese, korean and spanish including regional accents, though at the 4K tier audio narrows to chinese and english with everything else auto-translated to english. that is an odd limitation to hit late in a project.
the honest problem is value. it sits sixth and ninth on the artificial analysis text-to-video board while costing more than models above it. you are paying for the control surface and the 4K, not for output quality. one more thing: the widely-repeated claim that kling 3.0 runs at 60fps does not appear in kuaishou's own announcement — the 4K endpoint is confirmed, the frame rate is not.
pros
+only true native 4K video generation here
+per-shot control over duration, framing and camera movement
+five-language native audio with regional accents
+element consistency across a storyboarded sequence
cons
−ranks below cheaper models on quality leaderboards
−4K audio drops to chinese and english only
−the '60fps' claim is unverified — don't plan around it
best audio engineering, overtaken by its own stablemate
80/100
verdictstill the best-sounding model here and the only route to two and a half minutes of continuous output — but google's own omni flash beats it on quality at a quarter of the price.
best for
long-form assembly and enterprise work that needs indemnity and 48khz audio
price
$0.40 / second standard; Fast $0.10 at 720p; Lite $0.05
pricing note
4K is $0.60/second on standard; there is no veo 4 — anything you have read about one is speculation
free tier
yes
access
closed api (gemini, vertex) + flow
native audio
yes — 48khz
max duration
8s native, ~148s via extension
max resolution
4K
license
commercial, enterprise indemnity
veo 3.1's audio is genuinely the best engineered on this list: 48khz, with dialogue lip-synced properly, effects and ambience. if the output is going anywhere near a professional audio chain, that number matters.
video extension is its other real advantage. native generations are four, six or eight seconds, but chaining up to twenty extensions gets you to roughly 148 seconds of continuous video — far beyond anything else here, all of which cap out between ten and twenty seconds.
it does 4K, comes with google's enterprise trust and indemnity story, and integrates with flow.
the awkward part is that gemini omni flash — also google's — ranks first on the board where veo 3.1 now ranks tenth, at $0.10 a second against $0.40. if you are on veo 3.1 standard today and don't need the length or the 4K, you are paying four times over for a lower-ranked model. the fast tier at $0.10 is the more defensible choice.
pros
+48khz audio — the best sound quality in the category
+extension chains to roughly 148 seconds of continuous video
+native 4K at $0.60/second
+enterprise indemnity and google support
cons
−$0.40/second standard is 4× google's own better-ranked model
−only 4–8 seconds per native generation
−now ranks tenth on artificial analysis text-to-video with audio
verdictthird on both image-to-video boards at eight cents a second — a genuinely strong model attached to a genuinely serious reputational problem.
best for
cost-sensitive image-to-video where 720p is enough
price
$0.08 / second direct from xai
pricing note
fal charges $0.08 at 480p and $0.14 at 720p — about 75% more than xai direct at 720p
free tier
no
access
closed api (xai, fal)
native audio
yes — with lip sync
max duration
15 seconds
max resolution
720p
license
commercial via api terms
on the numbers this is one of the best-value models in the category: third on artificial analysis for image-to-video both with and without audio, beating models that cost three to four times more, with native synchronised audio including lip-synced dialogue.
buy it direct from xai rather than through fal — xai charges $0.08 a second flat, while fal charges $0.14 at 720p. that is a 75% markup for the same model.
the ceiling is 720p, with no 1080p option at all, and up to fifteen seconds.
the part that belongs in any honest review: xai faces multiple lawsuits over non-consensual sexualised deepfakes, reporting that the video tool produced sexualised imagery of a named public figure without being explicitly asked, a multistate attorney-general letter, and an amsterdam court threatening €100,000-a-day fines. safeguards added in january were reported to apply only to non-paying users. whatever you conclude about the model, putting it in a consumer-facing product is a decision to take with your legal team rather than your engineering one.
pros
+#3 on both artificial analysis image-to-video boards
+$0.08/second direct is excellent value
+native synchronised audio with lip sync
+fifteen-second clips
cons
−720p ceiling, no 1080p
−active deepfake litigation and regulatory pressure
verdictstatistically tied for fourth on image-to-video without audio, at nine cents a second for 1080p — if you don't need generated sound, this is the efficient frontier.
best for
high-volume 1080p b-roll where you're adding your own soundtrack anyway
price
$0.09 / second at 1080p without audio
pricing note
$0.115 with audio; cheaper tiers run from $0.025/second at 360p — unlike seedance, turning audio off genuinely saves money
free tier
yes
access
closed api (fal) + web app
native audio
yes — but the weak part
max duration
15 seconds
max resolution
1080p
license
commercial via api terms
pixverse v6 ranks fourth on artificial analysis for image-to-video without audio, effectively tied with grok 1.5, at $0.09 a second for 1080p. for silent footage that is the best quality-per-dollar on this list.
the pricing model is also more honest than most: audio is a per-second uplift you can decline. seedance charges you for audio whether you use it or not; here, switching it off actually reduces the bill.
it ships more than twenty cinema camera controls and a multi-shot engine, and handles one to fifteen seconds at up to 1080p.
the tell is in the leaderboard split: its with-audio ranking is markedly weaker than its without-audio ranking, which points at the audio component being the weak part. treat it as a strong silent-video model that happens to offer sound, rather than an audio-video model.
pros
+#4 on image-to-video without audio at $0.09/second for 1080p
+audio is optional and genuinely cheaper when off
+20+ cinema camera controls and a multi-shot engine
the best open-weight video model, with a licence that isn't what you think
73/100
verdicttwenty-second clips at up to 4K with genuine single-pass audio, running on your own hardware — just don't believe anyone who tells you it's apache 2.0.
best for
self-hosted pipelines needing long clips, 4K, and audio in one pass
price
free to self-host; $0.06 / second at 1080p on fal
pricing note
4K is $0.24/second hosted; the licence requires a paid commercial agreement above $10M annual revenue
free tier
yes
access
open weights + hosted api
native audio
yes — single pass
max duration
20 seconds
max resolution
4K
license
community licence, $10M revenue gate
ltx-2.3 is a 22-billion-parameter audio-video model from lightricks, and it is comfortably the strongest open-weight option here: leader among open weights on artificial analysis' image-to-video with audio board, twenty-second clips — the longest of any mainstream model — up to 4K, and 24 or 48fps.
it generates synchronised video and audio in a single forward pass, which most open models still can't do at all.
hosted on fal it is also cheap: $0.06 a second at 1080p, $0.24 at 4K, with a fast tier at two-thirds of that.
now the licence, because this is the single most misreported fact in the category — including on fal's own model pages, which state the model is apache 2.0. it is not. the actual licence file is the ltx-2 community license agreement, and it requires organisations with at least $10 million in annual revenue to obtain a paid commercial licence. it also prohibits use in products competing with lightricks' own, and specifies liquidated damages of double the licence fee for unauthorised commercial use. read the licence file, not the blog post.
pros
+best open-weight video model by a clear margin
+twenty-second clips — the longest here
+4K output with genuine single-pass audio
+cheap hosted option at $0.06/second for 1080p
cons
−not apache 2.0 — $10M revenue gate and anti-competition clause
−quality still well behind the closed frontier
−self-hosting a 22b video model needs serious gpu capacity
verdictthe base model no longer competes, but aleph video-to-video and act-two performance capture still have no real equivalent — you're buying the workflow, not the weights.
best for
editing workflows, performance capture, and video-to-video rather than raw generation
price
$0.12 / second (12 credits at $0.01)
pricing note
gen-3 alpha turbo and gen-4 aleph are being sunset on 30 july 2026 — check what your pipeline depends on
free tier
yes
access
closed api + web app
native audio
yes — on gen-4.5
max duration
not published
max resolution
720p native, 4K upscale
license
commercial on paid plans
runway's gen-4.5 generates natively at 720p, which in a field that reached 4K is a real limitation, and it appears in neither leaderboard's top ten. as a text-to-video model it has been passed.
what runway still has is everything around the model: aleph for video-to-video transformation, act-two for performance capture, a genuine editing timeline, and 4K upscaling. those are workflow tools that the model labs mostly haven't built, and for teams doing post rather than generation they remain the reason to be here.
runway has also become substantially a reseller. its pricing page now sells seedance 2.0, veo 3, veo 3.1, gemini omni flash and happy horse alongside its own model. that is a sensible read of where things are going, but it does mean the differentiator is the interface rather than the output.
one urgent operational note: gen-3 alpha turbo and gen-4 aleph are being sunset on 30 july 2026. if anything you run depends on those endpoints, that is days away, not months.
pros
+aleph video-to-video has no real equivalent elsewhere
+act-two performance capture and a real editing timeline
+aggregates seedance, veo, omni flash and happy horse in one place
+native audio on gen-4.5, generated in the same pass
cons
−native generation is only 720p
−absent from both leaderboards' top ten
−gen-3 alpha turbo and gen-4 aleph sunset 30 july 2026
the only one that speaks a professional colour pipeline
66/100
verdictnative 16-bit HDR and EXR export in ACES is unique here and genuinely matters to post houses — but it generates no audio in a year when everything else does.
best for
vfx and finishing work that has to survive a colour grade
price
~$1.20 per 5-second 1080p SDR generation
pricing note
credit-based and luma doesn't publish a per-credit rate; HDR is ~2× and HDR+EXR ~3×
free tier
yes
access
closed api + web app
native audio
no — third-party integration
max duration
20 seconds
max resolution
1080p, 16-bit HDR
license
commercial on paid plans
first, a correction: dream machine, ray2, ray3 and ray3.14 are all superseded. ray3.2 arrived in june and anything recommending luma dream machine is out of date.
luma's real differentiator is the colour pipeline. native 16-bit HDR and EXR export in ACES2065-1 is something no other model on this list offers, and for a vfx shop that has to composite and grade generated footage alongside real plates, that is the difference between usable and not.
it also offers up to sixteen keyframes per clip, which is the finest-grained motion direction available anywhere here, and modify video v2 handles up to twenty seconds at 1080p.
the gap is audio. luma generates none, integrating third-party tools instead, in a year when single-pass native audio became standard. pricing is also opaque — luma doesn't publish a per-credit dollar value, so the roughly $1.20 per five-second generation figure is an approximation rather than a quote.
pros
+native 16-bit HDR and EXR/ACES export — unique here
+up to sixteen keyframes for fine motion control
+twenty seconds at 1080p on modify video
+the right choice for professional finishing pipelines
verdicta genuinely capable model behind the least transparent pricing in the category — $1,000 minimum packages that expire monthly are hard to recommend to anyone small.
best for
teams already committed to minimax who can absorb the package minimums
price
opaque — sold in 'video points', from $1,000 per package
pricing note
minimax publishes no per-second rate; every per-second figure in circulation comes from resellers, and points expire after a month
free tier
no
access
closed api + consumer apps
native audio
unconfirmed in official docs
max duration
not published
max resolution
1080p
license
commercial via api terms
hailuo has a strong reputation for motion quality and a large consumer following, and the model itself is competitive.
the pricing is the problem and it is worth being blunt about it. minimax sells 'video points' in packages starting at $1,000 for 3,760 points, rising to $6,000 — and the points expire after a month. the documentation gives deduction examples rather than rates, so working out what a project costs means reverse-engineering it from worked examples.
every per-second figure you will find for hailuo comes from third-party resellers, not from minimax. we are not going to publish those as if they were the vendor's prices.
there are consumer plans from $9.99 and a developer token plan from $10 a month, which are more approachable, but the api story for anyone building at scale involves those package minimums. capable model, hostile commercial model.
pros
+strong motion quality and physics
+large consumer following and community knowledge
+consumer plans start at $9.99/month
+developer token plan from $10/month
cons
−no published per-second pricing at all
−api packages start at $1,000 and expire after a month
−native audio status not confirmed in official docs
verdictabsent from every leaderboard, and still the obvious pick if your legal team has watched what happened to bytedance this year.
best for
enterprises that need licensed training data and indemnification more than they need quality
price
$9.99 / month consumer; api from a ~$1,000 / month enterprise floor
pricing note
consumer plans do not include api access; firefly services is a separate enterprise contract
free tier
yes
access
subscription; api at enterprise tier
native audio
no — separate paid features
max duration
not published
max resolution
1080p
license
licensed data, indemnified
firefly video is not a competitive model on output quality and does not appear in any leaderboard we checked.
it is the only major video model with a clean commercial-safety story: trained on licensed and adobe stock content, with indemnification available. in a year when seedance drew cease-and-desists from disney and paramount and xai collected deepfake lawsuits, that is a product feature with real commercial value.
be precise about the audio claim, though. adobe offers audio translation, lip sync and sound-effect generation as separate credit-consuming features. that is not the same as single-pass native audio generation, and vendors are increasingly blurring the two.
the commercial structure is awkward for anyone small: consumer plans from $9.99 to $199.99 a month don't include api access at all, and the firefly services api starts around a thousand dollars a month. like the image model, firefly video increasingly routes to partner models rather than competing directly.
pros
+licensed training data with enterprise indemnification
+deep creative cloud and premiere integration
+the defensible choice when legal risk is the priority
+consumer plans start at $9.99/month
cons
−absent from every quality leaderboard
−api access starts around $1,000/month
−audio features are bolted on, not single-pass native
runs on one consumer gpu, banned in three jurisdictions
55/100
verdict8.3 billion parameters on a single rtx 4090 is a real achievement, and the licence excludes three major markets outright — check your geography before your gpu.
best for
fine-tuning research outside the eu, uk and south korea
price
free to self-host
pricing note
the licence expressly does not apply in the EU, the UK or South Korea — read it before you deploy
free tier
yes
access
open weights, self-hosted
native audio
no — separate foley model
max duration
~5 seconds
max resolution
720p
license
tencent community — excludes EU/UK/KR
hunyuanvideo 1.5 fits in 8.3 billion parameters and generates a clip in about seventy-five seconds on a single rtx 4090. for a research team or a solo developer without a cluster, that accessibility is the point.
it is still widely used as a fine-tuning base, and the weights are genuinely downloadable.
but the base model is silent — audio comes from a separate hunyuanvideo-foley model — and it tops out at roughly five-second clips at 720p. against ltx-2.3's twenty seconds at 4K with single-pass audio, it is a generation behind for open-weight work, and tencent has shipped nothing new on this line in 2026.
the licence deserves a careful read and is the reason it sits this low. it is the tencent hunyuan community licence, not an osi-approved open-source licence: organisations above 100 million monthly active users must request separate permission at tencent's discretion, outputs may not be used to train competing models, attribution is required, and — the clause most people miss — the licence expressly does not apply in the european union, the united kingdom or south korea. for a european company that is a hard blocker, not a footnote.
pros
+runs on a single consumer gpu
+genuinely downloadable weights
+still a solid fine-tuning base
+free to self-host
cons
−licence does not apply in the eu, uk or south korea
verdictsora 2 pro still ranks fifth on lmarena, and the api shuts down on 24 september 2026 — do not start anything here.
best for
nothing new — this is a migration notice, not a recommendation
price
~$0.10 / second standard, ~$0.30 pro
pricing note
openai's pricing page blocks automated access; these figures are secondary-sourced and largely academic now
free tier
no
access
closed api — retiring sept 2026
native audio
yes
max duration
not published
max resolution
1024p on pro
license
commercial until shutdown
this entry exists as a warning rather than a recommendation. openai is retiring sora in two stages: the consumer web and app experiences closed on 26 april 2026, and the api shuts down on 24 september 2026. developers on the videos api and sora 2 model aliases were formally notified in march.
the model itself is not the problem. sora 2 pro still sits fifth on lmarena, ahead of every veo variant, and its native audio was genuinely category-defining when it launched.
but a model with a published end-of-life two months out cannot be recommended for any new integration, and its low placement here reflects exactly that. capability is not the same as viability.
if you are already on it, treat migration as urgent. gemini omni flash is cheaper and ranks higher; seedance 2.0 and wan 2.7 are the closest matches for longer or higher-resolution work.
pros
+still #5 on lmarena text-to-video
+native audio was category-defining at launch
+strong physics and prompt adherence
cons
−api shuts down 24 september 2026
−consumer apps already closed in april
−no successor announced — openai has redirected the effort to research
verdictstill enjoyable for social-first content, but it appears in no leaderboard and has no real api — it has fallen out of the competitive tier.
best for
playful social formats and effects, not production or api work
price
free tier; paid from $8 / month
pricing note
commercial rights require a paid plan; there is no meaningful public developer api
free tier
yes
access
subscription consumer app
native audio
yes
max duration
not published
max resolution
1080p
license
commercial on paid plans
pika 2.5 handles physics interactions and synchronised audio, and its effects-driven approach is genuinely fun for social formats. that was enough to matter in 2024.
it is now absent from every leaderboard we consulted, and there is no meaningful public developer api — it is a subscription consumer app.
the free tier gives eighty credits a month and paid plans start at $8 with commercial rights, which makes it cheap to play with.
we would not build anything on it. a widely-referenced 'pika 3.0' appears to be projected rather than shipped, so we're not ranking on it.
hands-on prompting around the things that actually break: a physics shot, a dialogue shot to check lip sync, a camera move, and a pair where the same character has to survive a cut. qualitative, and used to sanity-check the leaderboards rather than to produce a score of our own.
cost per second is the vendor's published rate at the resolution named in the entry, taken from the vendor's pricing page or the fal model page on the review date. video pricing is tiered by resolution and often by whether audio is on — we state which tier the number refers to, because a single blended 'price per second' for this category would be fiction.
native audio means generated in the same pass as the video. a model that hands off to a separate sound-effects or lip-sync product is marked as not having it, however good the combined result is.
we read the actual licence file. three of the most widely-cited 'apache 2.0' claims in video are wrong — ltx-2 ships a community licence with a $10m revenue gate and an anti-competition clause, wan's public weights stop at 2.2, and tencent's hunyuan licence explicitly does not apply in the eu, the uk or south korea.
we cross-check against the artificial analysis and lmarena video leaderboards. both currently agree on the top two, which is a stronger signal than either alone.