R4. AI video models and pricing, as of 2026-09-04

Source of truth: fal's model API queried with the project key on 2026-09-04 (154 text-to-video, 233 image-to-video, 223 video-to-video endpoints), plus fal model pages, lab blogs, and X posts fetched the same day by a research agent. Prices are fal page text as of today. Promo prices expire: H3 Max Sep 7, H3 Max Director Sep 14.

Flagship models

Lab / model Release fal endpoint fal price per second Max clip Max res Native audio Ref-to-video Multi-shot Source
Alibaba Wan 3.0 Prime GA Aug 24 2026 alibaba/wan-3.0-prime/text-to-video, /image-to-video, /reference-to-video; also alibaba/wan-3.0/* $0.068 480p, $0.14 720p, $0.28 1080p 2 to 30s single pass 1080p Yes, same pass 10 img + 5 video + 5 audio + doc/webpage Long take Reuters, fal
ByteDance Seedance 2.5 Jul 31 2026; fal Aug 7 bytedance/seedance-2.5/text-to-video, /image-to-video, /reference-to-video ~$0.22 480p, ~$0.47 720p, ~$1.16 1080p (token billed, $0.0214 per 1k) 4 to 30s single pass 1080p (720p recommended) Yes, joint latent Up to 50 refs (30 img, 10 video, 10 audio) Yes, single pass Seed blog, fal
MiniMax H3 Max (fal post-trained) Aug 26 2026; PR Sep 1 minimax/h3-max/text-to-video, /image-to-video, /reference-to-video, /director; minimax/h3-max-turbo/* $0.0125 480p, $0.02 768p promo to Sep 7 (list $0.05 / $0.08); Turbo $0.01 768p; Director $0.02 promo, 60s minimum 5 to 15s 768p on fal Yes, stereo Yes Director endpoint MiniMax docs, fal, PR
MiniMax H3 (open weights) Jul 30 2026 minimax/h3/text-to-video, /image-to-video, /reference-to-video, /text-to-video/lora, /image-to-video/lora higher than H3 Max per X reports ($1.30 vs $0.20 per clip) 4 to 15s 2K Yes Yes No kfgo
Google Gemini Omni 1.1 Flash Omni Flash May 19 2026; 1.1 Aug 27 2026 google/gemini-omni-flash/v1.1/text-to-video, /image-to-video, /reference-to-video, edit $0.03 360p, $0.10 720p, $0.15 1080p, $0.30 4K 8s default; extend to 40s per Google 4K Yes Yes Via extend Google blog, fal
Google Veo 3.1 (Lite tier Apr 2026; no Veo 4 found) Oct 2025 fal-ai/veo3.1, /fast, /lite, /reference-to-video, /image-to-video, /first-last-frame-to-video, extend $0.20 silent, $0.40 audio (720/1080); 4K $0.40 / $0.60; fast $0.10 / $0.15 8s; extend to ~148s 4K Yes Yes Via extend fal, Vertex Lite
Kuaishou Kling 3.0 / 3.0 Omni (O3); 3.0 Turbo Jun 17; native 4K Apr 23 Feb 5 2026 fal-ai/kling-video/v3/{standard,pro}/*, o3/*, v3/turbo/*, v3/4k/* $0.112 silent, $0.168 audio, $0.196 with voice control 3 to 15s 1080p (4K request types exist) Yes (EN/ZH voice) Elements on i2v / o3 multi_prompt 1 to 6 shots within 15s Kuaishou PR, fal
Lightricks LTX 2.5 (open weights) Aug 11 2026 lightricks/ltx-2.5/text-to-video/pro, /fast, i2v $0.12 720p, $0.17 1080p 6/8/10s pro; up to 20s in other variants 1080p Yes No r2v endpoint seen Claimed VentureBeat, fal
BFL FLUX 3 video Jul 23 2026 early access; fal Aug 4 blackforestlabs/flux-3/text-to-video, /image-to-video, /keyframes-to-video, /first-last-frame-to-video, /draft variants $0.17 720p, $0.29 1080p 20s 1080p Yes, multilingual dialogue Image refs via i2v only Agentic chaining BFL blog, fal X
xAI Grok Imagine Video 1.5 GA Jun 17 2026 xai/grok-imagine-video/v1.5/text-to-video, /image-to-video, /reference-to-video, extend $0.08 480p, $0.14 720p, $0.25 1080p 6s 1080p Yes Yes No x.ai, fal
Luma Ray 3.2 Jun 9 2026 luma/agent/ray/v3.2/text-to-video, /image-to-video not captured 20s 1080p Yes 16 keyframes (Luma app) Keyframes Luma
Runway Gen-4.5 Dec 1 2025 not on fal n/a 10s 1080p HDR Yes Yes No Runway
OpenAI Sora 2 Sep 30 2025; product shut down Apr 26 2026 fal-ai/sora-2/* all deprecated n/a n/a n/a n/a n/a n/a OpenAI
Vidu Q3 (Shengshu) Jan 30 2026; R2V Apr 13 fal-ai/vidu/q3/* not captured 16s not stated Yes Yes Smart cuts PRNewswire
Pixverse v6 / C1 Mar 29 / Apr 8 2026 fal-ai/pixverse/v6/*, c1/* not captured not captured not captured Yes (switch) C1 reference-to-video multi-clip switch fal catalog

Talking-head and lip-sync specialists on fal

Image models for reference sheets and keyframes

All active in the fal catalog on 2026-09-04: fal-ai/nano-banana-pro (+ /edit), fal-ai/nano-banana-2 (+ /edit), fal-ai/bytedance/seedream/v5/lite/text-to-image and /edit, fal-ai/bytedance/seedream/v4.5/*, fal-ai/gpt-image-1.5, fal-ai/flux-2-pro (+ /edit), fal-ai/gemini-3-pro-image-preview (+ /edit), fal-ai/z-image/*. Pricing not yet captured; add in Phase 3.

Community read on realistic talking humans, Aug 7 to Sep 4 2026

Single-creator tests, not reproduced. Ranking flips with content type and safety filters.

Net: Seedance 2.5 leads on face realism and 30s multi-shot consistency. Wan 3.0 and H3 / H3 Max lead on dialogue lip-sync per dollar. Kling 3.0 audio draws repeated complaints. Veo 3.1 is cited as strong but is the oldest and most expensive per second. Sora is gone.

Bake-off plan (Phase 5)

Same three shots on each: (1) solo confessional, 8s, one line of dialogue; (2) two-person scheme in the control room, 10s, two lines; (3) implied-action montage beat, 6s, no dialogue.

  1. Wan 3.0 Prime reference-to-video, 720p draft then 1080p.
  2. Seedance 2.5 reference-to-video, 720p only unless it wins outright.
  3. MiniMax H3 Max reference-to-video and image-to-video, 768p (draft workhorse regardless of outcome). Not director: R5 found it is a WebRTC realtime stream with a 60 s minimum bill, not a batch multi-shot endpoint. Batch multi-shot on H3 Max is "CUT." inside the prompt.
  4. Gemini Omni 1.1 Flash reference-to-video, 1080p, shot 1 only (hero confessional fallback).

Score per clip: identity match to reference sheet, lip-sync, acting naturalness, reality-TV look, filter refusals, cost. Winner per shot type, not overall.

Luxin overlap

Luxin's hosted runtime (registry 2026-08-12) wires 18 video models: Veo 3.1 fast/lite, Kling v3 standard, Seedance 2.0 fast, LTX 2.3 fast, Wan 2.7, Pixverse v6, Hailuo-02, Ray 3.2, ltx-video-13b. None of Wan 3.0, H3 Max, Omni Flash 1.1, Seedance 2.5 are wired. This project calls fal directly.

Not verified

Gemini Omni Flash duration options beyond 8s on fal. Kling v3 fal output resolution enum. alibaba/wan-3.0 non-Prime price. Vidu Q3 max resolution. Luma, Pixverse, Vidu fal prices. Image model prices. All X-post quality claims. Whether Seedance 2.5 filters reject "producer hands contestant a drink" style prompts (test in Phase 5).