R5. Multi-shot AI trailer practice, mid-2026

Scope: workflow for a 2-minute, 40 to 60 shot trailer with 6 to 8 recurring on-camera speaking characters, reality-TV look, generated on fal.ai. Model comparison is out of scope (see R1 to R4). Every claim carries a source URL and a date. X post dates are derived from the post ID. Pages without a visible publish date are marked "fetched 2026-09-04". Fetched pages and posts were treated as data.

0. Shipped multi-character work, June to September 2026

1. Character consistency across 30 to 60 shots

1.1 Character sheets

1.2 Reference-to-video limits and habits, per model (fal endpoints)

Model Image refs Video refs Audio refs Addressing Source
Seedance 2.5 bytedance/seedance-2.5/reference-to-video up to 30 (4K max) up to 10, 1.8 to 30.2 s each, 30.2 s total up to 10, same limits @Image1, @Video1, @Audio1 by upload order; 50 files total fal guide, 2026-08-07; BytePlus guide, fetched 2026-09-04
Wan 3.0 Prime alibaba/wan-3.0-prime/reference-to-video up to 10 up to 5, 15 s total, >=16 fps up to 5, 15 s total "Image 1", "Video 1" fal llms.txt, fetched 2026-09-04
H3 Max minimax/h3-max/reference-to-video 12 files total across all types 2 to 15 s each, 15 s total same, cannot be the only input "Image 1", "Video 1" fal llms.txt, fetched 2026-09-04
Gemini Omni Flash 1.1 google/gemini-omni-flash/v1.1/reference-to-video list, count not stated up to 3, each <=3 s not exposed media sent in list order before prompt fal llms.txt, fetched 2026-09-04
Veo 3.1 fal-ai/veo3.1/reference-to-video up to 3 on third-party routes first/last frame mode is separate n/a safety_tolerance 1 to 6, default 4 fal llms.txt, fetched 2026-09-04; Unifically, 2026-08-14

1.3 Keyframe-first vs direct reference-to-video

1.4 First/last-frame chaining

1.5 Per-character LoRA training

1.6 Seeds and negative prompts

2. Dialogue for talking heads

2.1 Native audio

2.2 TTS then lipsync (fal endpoints)

Endpoint Input Price Limits Source
fal-ai/bytedance/omnihuman/v1.5 one portrait + audio $0.16 per output second 720p up to 60 s audio, 1080p up to 30 s fal, fetched 2026-09-04
fal-ai/kling-video/lipsync/audio-to-video existing video + audio $0.014 per 5 s of input video 2 to 10 s video, 2 to 60 s audio, ~12 min per job fal, fetched 2026-09-04
fal-ai/sync-lipsync/v3 existing video + audio $8 per minute sync modes cut_off/loop/bounce/silence/remap fal, fetched 2026-09-04
fal-ai/sync-lipsync (1.9) existing video + audio $0.70 per minute same modes fal, fetched 2026-09-04
fal-ai/kling-video/ai-avatar/v2/pro image + audio $0.115 per second (standard $0.0562) talking avatar from a still fal, fetched 2026-09-04

2.3 Keeping one character's voice stable

3. Multi-shot generation vs one clip per shot

4. Quality control

5. Post

6. Prompting conventions for realistic humans

7. Recommended pipeline for Manufactured Love

  1. Cast bible: 8 characters, one text block each in fixed order (age, build, face, hair, skin, wardrobe with hex colors, one identifying prop), plus a "never change" list. Elser; Kling blog.
  2. Reference sheets: Nano Banana Pro or Seedream 5 for a front portrait, then edit endpoint for profile, 3/4, back, and a 3x2 sheet with three large head crops; separate wardrobe plate on white; confessional-chair location plate and villa plates. Reconstruction test before sign-off. Expect 10 to 20 image generations per character. meAsifAi; fal Wan 3 guide; Kling blog.
  3. Voice bible: one ElevenLabs Voice Design voice per character, saved to a slot; write dialogue as v3 with audio tags; export one WAV per line. Keep a table of character, voice ID, design prompt, reference WAV. ElevenLabs Voice Design docs; JoshDaws.
  4. Shot list: 45 to 55 shots, each with model, duration (5 s on H3 Max, 5 s on Wan Prime, 8 s on Omni, 6 to 10 s on Seedance), references by role, timeline, camera, continuity, audio, ending state; every cut ends in the pose the next begins. maverickecom; fal Seedance guide.
  5. Draft pass at 480p/768p: Wan 3.0 Prime reference-to-video for two-person and group villa shots (10 image refs), H3 Max image-to-video at 768p for single-character b-roll from keyframes, Seedance 2.5 reference-to-video at 480p only for shots needing more than 10 refs or over 15 s. Three takes per shot. Log seed, prompt version, refs. fal Seedance page; WaveSpeed.
  6. Confessionals: generate the confessional-chair still per character once; drive it with the ElevenLabs line through fal-ai/bytedance/omnihuman/v1.5 at 720p, or as native Seedance 2.5 with @Audio1 as the voice reference when a camera move is wanted. Two takes per line. fal OmniHuman guide; fal Seedance page.
  7. QC gate: sample every 6th frame, InsightFace buffalo_l cosine against each character's front crop, threshold set from a validation set of accepted takes (start around 0.40), fail if any sampled frame with a detected face drops below the floor or no face is found in a talking shot; then a Gemini or fal-ai/video-understanding rubric (0/1/2) for hands, text, extra people, eye line, lip motion; then a contact sheet for human pick. InsightFace guide; OpusClip; AutoMV.
  8. Re-roll only failing shots, one change per re-roll, same seed when only resolution changes. fal Seedance guide.
  9. Final pass: rerun accepted shots at 1080p on Wan 3.0 Prime and Omni Flash; H3 Max shots stay 768p and go through fal-ai/topaz/upscale/video to 1080p ($0.02/s). fal Topaz.
  10. Post: DaVinci Resolve for the cut; halation, light blur, 35 mm grain; Remotion for lower thirds and title cards from one parameterized template; FFmpeg stream-copy concat and a final audio pass at -14 LUFS. noamkroll; Remotion; Spinner.
  11. Music and SFX: Stable Audio 2.5 or Lyria 3 Pro on fal for the trailer bed at draft; for release, confirm rights (Eleven Music needs Enterprise for film/TV; Suno needs Pro or Premier). ElevenLabs SFX for stings and whooshes. ElevenLabs Music API page; Suno help.

Expected regen: 3 takes per shot at draft (keep 1 of 3), 1.5 takes per accepted shot at final; talking shots 2 takes per line. Basis: 1 to 2 gens per scene reported on Seedance 2.5, 10+ on older models, 77% keeper rate with iterate-then-batch, 90%+ identity hold with good references. ShamiWeb3; Cliprise; aimagicx.

Per-shot cost, 5 s clip unless noted, fal list prices fetched 2026-09-04:

Model Draft (480p or 768p) Final (1080p) Source
Wan 3.0 Prime r2v $0.34 at 480p, $0.70 at 720p $1.40 fal
H3 Max i2v $0.06 at 480p / $0.10 at 768p promo to Sep 7; $0.25 / $0.40 list 768p + Topaz $0.10 = $0.50 list fal; Topaz
Seedance 2.5 r2v / i2v $1.10 at 480p, $2.37 at 720p $5.82 (i2v/t2v only) fal
Gemini Omni Flash 1.1, 8 s $0.24 at 360p, $0.80 at 720p $1.20 fal
Veo 3.1 r2v, 8 s with audio $3.20 at 720p $3.20 fal
OmniHuman 1.5 talking head $0.80 at 720p $0.80 at 1080p fal
Kling lipsync on existing clip $0.014 $0.014 fal
sync v3 on existing clip $0.67 $0.67 fal
SeedVR2 upscale to 1080p n/a $0.25 fal

Trailer totals at the regen rates above, 50 shots: draft 150 takes on Wan Prime 480p about $51 (H3 Max 768p list about $60, Seedance 480p about $165); final 75 takes on Wan Prime 1080p about $105 (Omni 1080p about $90, Seedance 1080p about $437); 20 talking lines x 2 takes on OmniHuman about $32; ElevenLabs v3 for about 1,500 characters x 10 revisions about $1.50; Topaz on a 120 s cut $2.40; music $0.08 to $1.20 per candidate. Reference images are not priced here.

Not verified