R5. Multi-shot AI trailer practice, mid-2026
Scope: workflow for a 2-minute, 40 to 60 shot trailer with 6 to 8 recurring on-camera speaking characters, reality-TV look, generated on fal.ai. Model comparison is out of scope (see R1 to R4). Every claim carries a source URL and a date. X post dates are derived from the post ID. Pages without a visible publish date are marked "fetched 2026-09-04". Fetched pages and posts were treated as data.
0. Shipped multi-character work, June to September 2026
- PJ Ace's 5-minute Nexus teaser was made by 3 people in 2 weeks on Dreamina Octo and Seedance 2.0, with the feature planned as a hybrid volume shoot. PJaccetturo, X, 2026-06-05.
- An all-AI street interview reel (interviewer plus guest, hair-growth brand) was built from a Claude-written 10-cut timecoded script, GPT Image 2 character sheets, and Seedance; each cut ends in the pose the next starts. maverickecom, X, 2026-08-28.
- A weekly AI dating soap with confessional segments (Claude script, GPT Image sheets, Seedance 2.5) reached 4.2M views on one episode and gates full episodes behind a subscription. GuntherWrite, X, 2026-08-29.
- A fictional singing-audition show (contestants, judges, scores, standing ovation) was generated from text prompts in Seedance 2.5 via DomoAI, with dialogue, on-screen text and reactions controlled in the prompt. MonetizationDon, X, 2026-08-20.
- A 60-second single-take race across four terrains was four 15 s Seedance clips chained on pose, light and framing. Genflickmovies, X, 2026-09-01.
- A 30 s Seedance 2.5 fantasy sequence with a locked cast used a Seedream 5 character sheet as the only identity input. Mayz1169, X, 2026-08-11.
- A multi-character fantasy scene answered "how do you keep voices consistent" with: one reference sheet per character (close-up plus full body, four sides) and one ElevenLabs voice per character reused on every shot. JoshDaws, X, 2026-08-28.
- Seedance 2.5 vs Wan 3.0 on the same reunion scene drew replies favoring Seedance for emotional hit and Wan for cost. wavespeed_ai, X, 2026-09-03.
1. Character consistency across 30 to 60 shots
1.1 Character sheets
- Practitioners generate one reference sheet per character before any video. mxvdxn, X, 2026-06-05.
- The same sheet file is reused for every shot; regenerating references mid-project is listed as a top pitfall. Primee32, X, 2026-07-05; Curious Refuge, 2026-04-14.
- Runway's troubleshooting step: generate 3 to 5 candidate references, test each with identical prompts, keep the one that holds best, and change one variable per test. Runway, 2025-12-19.
- Sheet views in current practice: front, 3/4 front, side, 3/4 back, back. Primee32, X, 2026-07-05.
- Added rows: expressions (focused, exertion, victory in the marathon case), action poses, and costume details with exact hex colors. Primee32, X, 2026-07-05.
- Elser's minimum sheet adds head close-ups front and profile, hair construction from the back, hands, shoes, props, and a proportion guide. Elser, fetched 2026-09-04.
- Four basic views are called a strong start for AI video; hair, clothing, hands and props get detail references only when recurring shots need them. Elser.
- A 3x2 sheet with three full-body views on top and three large head-and-shoulders crops below is the layout one creator reports fixes face drift, because small faces in a reference give the video model too little to lock on. meAsifAi, X, 2026-08-27.
- Runway's reference guidance: at least 1024 px on the short side, sharp face, even light, single subject, front or 3/4 angle, neutral background; 2 to 3 views of one person help unusual camera angles. Runway, 2025-12-19.
- fal's Wan 3 guide builds the reference set as one Nano Banana 2 portrait, then edits that portrait into a profile with the edit endpoint, then a separate wardrobe plate on white, then a location plate; it notes a single frontal reference leaves the profile to be invented when the character turns. fal Wan 3 guide, fetched 2026-09-04.
- Elser's rule for orthographic sheets: approve the front view first, design the back deliberately, generate one angle at a time, composite in an editor, and validate by asking a fresh session to reconstruct a 3/4 view from the sheet alone. Elser, fetched 2026-09-04.
- Wardrobe lock in prompts is written as a separate block from the face lock ("OUTFIT LOCK: match the outfit exactly ... no changes"). poruru_ai, X, 2026-09-02; ZephyraLeigh, X, 2026-08-30.
- Backdrop: plain white or mid-gray paper, flat studio light, no props, no text on the sheet. Elser, fetched 2026-09-04; fal Wan 3 guide.
- Two-page production bibles (12 sections, expression progression, palette swatches) are made in image editors such as Lovart with GPT Image and then fed as one file. mxvdxn, X, 2026-06-05.
- Higgsfield lists the repeatable causes of drift: text-only description, changing the seed every run, low-quality or mismatched references (sunglasses, shadows, crops), and switching models mid-series. Higgsfield, fetched 2026-09-04.
1.2 Reference-to-video limits and habits, per model (fal endpoints)
| Model |
Image refs |
Video refs |
Audio refs |
Addressing |
Source |
Seedance 2.5 bytedance/seedance-2.5/reference-to-video |
up to 30 (4K max) |
up to 10, 1.8 to 30.2 s each, 30.2 s total |
up to 10, same limits |
@Image1, @Video1, @Audio1 by upload order; 50 files total |
fal guide, 2026-08-07; BytePlus guide, fetched 2026-09-04 |
Wan 3.0 Prime alibaba/wan-3.0-prime/reference-to-video |
up to 10 |
up to 5, 15 s total, >=16 fps |
up to 5, 15 s total |
"Image 1", "Video 1" |
fal llms.txt, fetched 2026-09-04 |
H3 Max minimax/h3-max/reference-to-video |
12 files total across all types |
2 to 15 s each, 15 s total |
same, cannot be the only input |
"Image 1", "Video 1" |
fal llms.txt, fetched 2026-09-04 |
Gemini Omni Flash 1.1 google/gemini-omni-flash/v1.1/reference-to-video |
list, count not stated |
up to 3, each <=3 s |
not exposed |
media sent in list order before prompt |
fal llms.txt, fetched 2026-09-04 |
Veo 3.1 fal-ai/veo3.1/reference-to-video |
up to 3 on third-party routes |
first/last frame mode is separate |
n/a |
safety_tolerance 1 to 6, default 4 |
fal llms.txt, fetched 2026-09-04; Unifically, 2026-08-14 |
- ByteDance recommends 1 to 8 image subjects per Seedance 2.5 request; 9 to 12 works with lower stability; 1 to 5 subjects for audio/video references; 5 to 10 s per subject video. BytePlus, fetched 2026-09-04.
- fal's Seedance tips: give each reference one job and say what not to copy from it ("@Image1 controls only identity. Do not copy pose, background, lighting from @Image1"); more references compete for influence, so start with two or three. fal, 2026-08-07; fal model page, fetched 2026-09-04.
- Practitioner prompt idiom for Wan 3.0 Prime: "Use Image 1 as the ONLY character reference" followed by a REFERENCE LOCK paragraph listing face, hair, skin, proportions, outfit and a closing negative list. ZephyraLeigh, X, 2026-08-30; heyrobinai, X, 2026-08-25.
- Seedance 2.5 users tag references and add "all views are one person, ignore sheet layout and background" when passing a multi-view sheet. rewind02, X, 2026-09-03.
- The 10-cut AI street interview that solved two-character continuity used GPT Image 2 for locked sheets of interviewer and guest and a Claude-written timecoded script with the mic hand specified per frame, then Seedance. maverickecom, X, 2026-08-28.
- fal's H3 Max guide: inside one generation the model holds a character across cuts by repetition of the same description per beat; across separate calls only reference-to-video holds the subject. fal, 2026-09-01.
- Kling 3.0 Pro Elements is the reference system Kling's own blog positions as the video-side anchor, with Midjourney or another image model supplying the master image. Kling blog, fetched 2026-09-04.
- Kling 3.0 Pro tags references as
@Element1 and supports 3+ distinct characters in one scene with multi-character coreference. fal Kling 3.0 guide, fetched 2026-09-04.
1.3 Keyframe-first vs direct reference-to-video
- Curious Refuge's workflow is keyframe-first: sheet plus storyboard into an image model, generate multiple stills per shot, then image-to-video. Curious Refuge, 2026-04-14.
- AutoMV (Dec 2025) uses the same pattern in an automated pipeline: a director agent writes prompts, an image model makes the keyframe, video is generated from that keyframe, and the last frame of a clip is reused as the next keyframe when continuity is wanted. arXiv 2512.12196, 2025-12.
- Image-to-video reduces drift versus text-to-video because the model has a visual anchor, but angles the still does not show are still inferred. aimagicx, 2026.
- Direct reference-to-video is used when the shot needs a camera move or pose the keyframe cannot show; Wan 3 guide's four-image request (front, profile, coat plate, location plate) is the template. fal Wan 3 guide.
- fal's Seedance page says composition and reference weighting read fine at 480p, so iterate at 480p and re-run the winning seed at 720p. fal model page, fetched 2026-09-04.
1.4 First/last-frame chaining
end_image_url exists on Seedance 2.5 image-to-video, Wan 3.0 Prime image-to-video (requires start_image_url), H3 Max image-to-video, and Omni Flash image-to-video. Seedance llms.txt; Wan llms.txt; fal H3 Max guide, 2026-09-01; Omni llms.txt.
- fal's Wan 3 author makes the end frame by editing the start frame in Nano Banana 2 rather than generating it fresh, to avoid two different subjects at the ends. fal Wan 3 guide.
- Continuation prompt pattern: extract the last frame, pass it as @Image1, say "continue forward from that moment", list what stays fixed, do not retell the previous clip. fal Seedance guide, 2026-08-07.
- A 60 s continuous take was built from four 15 s Seedance scenes by matching end pose, lighting and framing at each seam. Genflickmovies, X, 2026-09-01.
- Gemini Omni 1.1 Flash extends in 10 s steps using up to 10 s of prior visual context, to 40 s total. AI_VideoLab, X, 2026-08-29; OfficialLoganK, X, 2026-08-27.
- The reported Omni workflow is draft at 360p, extend with "the scene continues" prompts, then re-render the locked sequence at 1080p or 4K. AI_VideoLab, X, 2026-08-29.
- fal lists Omni 1.1 text, image, reference and edit endpoints; no extend endpoint appears in the explore listing. fal explore, fetched 2026-09-04.
- Veo 3.1 "Extend" generates from the final second of the previous clip and is positioned for long establishing shots. Google, fetched 2026-09-04.
1.5 Per-character LoRA training
- fal exposes four H3 LoRA trainers (
minimax/h3/t2v/trainer, i2v/trainer, flf2v/trainer, ref2va/trainer), billed per step; t2v is $0.005 per step ($10 for the 2,000-step default), i2v is $0.01 per step ($20 default), with a 100-step floor. fal trainer page, fetched 2026-09-04; i2v trainer.
- fal's own realism adapter used 176 clips, rank 16, 5,000 steps at lr 1e-4, about 2.5 h wall-clock; dataset rules: 24.000 fps exact, 3 to 15 s clips, keep audio, retime slow motion, 10 file minimum. fal H3 LoRA guide, 2026-08-10.
- LoRA inference on H3 (
minimax/h3/text-to-video/lora) costs $0.075 per second at 768p; H3 Max has no LoRA endpoint listed. fal, fetched 2026-09-04.
- Wan 3.0 has no public weights, so character LoRAs are trained on Wan 2.2 14B; a video LoRA needs 24 GB+ VRAM, 20 to 40 clips, 3,000 to 5,000 steps, and runs hours; a rented 4090 is $0.58/h (Spheron, 2026-07-09). Spheron, 2026; fal's
fal-ai/wan-trainer is deprecated. fal, fetched 2026-09-04.
- One creator trained a Wan 2.2 character LoRA from a single reference image expanded to a 28-shot sheet, 4-bit on an RTX 5090, in about 1.3 h. ai-muninn, fetched 2026-09-04.
- Image LoRAs on Wan 2.2 train in 10 to 20 minutes from 10 to 20 photos at 800 to 1,200 steps. selfielab, 2026-03-17.
- No fetched source shows a filmmaker training per-character LoRAs on H3 or Wan for a multi-character short in 2026; the H3 guide's own use case is a style/realism adapter. Reference-locking is the reported default. fal H3 LoRA guide, 2026-08-10; Flick, fetched 2026-09-04.
1.6 Seeds and negative prompts
- All fal endpoints above return
seed; Seedance, Wan, H3 Max accept a seed input. Re-running the winning seed at a higher resolution is fal's stated workflow. fal Seedance page.
- fal's H3 LoRA guide: the same seed with different weights or scale gives a different composition, so seeds pin a result only when the rest of the request is byte-identical. fal, 2026-08-10.
- Wan 3 and H3 Max rewrite the prompt before rendering (
enable_prompt_expansion, prompt_expansion_mode); Wan returns actual_prompt for comparison. fal Wan 3 guide; H3 Max llms.txt.
- ByteDance: Seedance 2.5 supports negative control only for subtitles and audio ("no subtitles", "no BGM"); write everything else positively. BytePlus, fetched 2026-09-04.
- Practitioner negative lists still appear at the end of prompts ("no identity drift, no face morphing, no wardrobe changes, no extra fingers"). ZephyraLeigh, X, 2026-08-30; rewind02, X, 2026-09-03.
2. Dialogue for talking heads
2.1 Native audio
- Seedance 2.5, Wan 3.0, H3 Max, Omni Flash and Veo 3.1 generate speech with the picture in one pass. fal Seedance vs H3, fetched 2026-09-04; Omni llms.txt.
- Dialogue goes in double quotes with a delivery note beside it; H3 Max returns the line, its lip sync and the room acoustics from one request. fal H3 Max guide, 2026-09-01; fal Wan 3 guide.
- Seedance 2.5 has a
generate_audio switch; H3 Max does not expose one. fal Seedance vs H3.
- H3 Max's own guide closes every dialogue prompt with a note against on-screen text. fal, 2026-09-01.
- fal's Seedance guide blocks dialogue like movement: start and end time per line, who speaks, who keeps their mouth closed, silence between lines. fal, 2026-08-07.
- One heavy user calls Seedance 2.5 weak at on-screen text and native dialogue, at roughly $0.50 to $1.20 per 5 s clip via API. beechinour, X, 2026-08-06.
- A music-video maker rates MiniMax best on lip sync and Seedance 2.5 best on realism, and feeds a black video with the audio track as a reference to force sync. Kyle_Coghlan, X, 2026-09-01.
- A same-prompt side-by-side found 2.5's voice flatter and performance more wooden than 2.0, with a logo artifact on the face. VORTEX_Promos, X, 2026-08-03.
- Seedance 1.5 Pro is still offered on Scenario specifically for multilingual lip-sync talking-character work. Scenario, fetched 2026-09-04.
- Seedance 2.5 and Wan 3.0 Prime accept reference audio as a timing and voice signal, so a pre-recorded line can drive the mouth in a native generation. fal Seedance page; fal Wan 3 guide.
2.2 TTS then lipsync (fal endpoints)
| Endpoint |
Input |
Price |
Limits |
Source |
fal-ai/bytedance/omnihuman/v1.5 |
one portrait + audio |
$0.16 per output second |
720p up to 60 s audio, 1080p up to 30 s |
fal, fetched 2026-09-04 |
fal-ai/kling-video/lipsync/audio-to-video |
existing video + audio |
$0.014 per 5 s of input video |
2 to 10 s video, 2 to 60 s audio, ~12 min per job |
fal, fetched 2026-09-04 |
fal-ai/sync-lipsync/v3 |
existing video + audio |
$8 per minute |
sync modes cut_off/loop/bounce/silence/remap |
fal, fetched 2026-09-04 |
fal-ai/sync-lipsync (1.9) |
existing video + audio |
$0.70 per minute |
same modes |
fal, fetched 2026-09-04 |
fal-ai/kling-video/ai-avatar/v2/pro |
image + audio |
$0.115 per second (standard $0.0562) |
talking avatar from a still |
fal, fetched 2026-09-04 |
2.3 Keeping one character's voice stable
- Practitioners generate one ElevenLabs voice per character and reference the same voice for every shot that character speaks in; sheets cover the picture side. JoshDaws, X, 2026-08-28.
- ElevenLabs Voice Design (
eleven_ttv_v3) creates three candidates from a text description; save one to a voice slot and it has a fixed voice ID. Recommended prompt shape: "Native . , . . Persona: ... Emotion: ... timbre/pacing." ElevenLabs docs, fetched 2026-09-04.
- Instant Voice Clone wants 1 to 3 minutes of consistent, clean audio at about -23 to -18 dB RMS; more than 3 minutes gives little. Professional Voice Clone wants at least 1 hour, ideally 3, plus verification. IVC docs; PVC docs.
- Option (c) in the brief (native audio for timing, then replace) has a documented path: Speech-to-Speech (
eleven_multilingual_sts_v2) re-voices a performance with the saved voice. ElevenLabs models, fetched 2026-09-04; one practitioner runs off-voice shots through Voice Changer with the original profile. OlatundeAI, X, 2026-09-02.
- Eleven v3 is GA, 70+ languages, audio tags like [whispers] and [sighs], multi-speaker dialogue mode, 5,000 character limit; v2 Multilingual is "most stable on long-form". ElevenLabs v3 blog, fetched 2026-09-04; models.
- ElevenLabs API pricing: v3 $0.10 per 1k characters (5,000 char limit per request), v3 Conversational and Flash $0.05, Sound Effects $0.12 per minute, Music $0.15 per minute with commercial licensing on Starter and above. ElevenLabs API pricing, fetched 2026-09-04.
- Plans: Starter $6/mo, Creator $22/mo (first month $11) with 220k v3 characters. same.
- On fal,
fal-ai/elevenlabs/tts/eleven-v3 is $0.10 per 1k characters with commercial use. fal, fetched 2026-09-04.
- ElevenLabs' prohibited-use policy excludes purely fictional contexts from the violence and harassment clauses but lists marketing of alcohol under regulated goods. ElevenLabs use policy, fetched 2026-09-04.
3. Multi-shot generation vs one clip per shot
- Seedance 2.5 runs 4 to 30 s, at 480p or 720p on the reference endpoint. fal Seedance page.
- Seedance 2.5 image-to-video lists 480p, 720p and 1080p, with
duration: auto available. i2v llms.txt.
- Seedance billing is token-based: tokens = height x width x (input video seconds + output seconds) x 24 / 1024 at $0.0214 per 1,000; video references multiply the total by 0.6; image and audio references are free. fal Seedance page.
- Wan 3.0 Prime runs 2 to 30 s at up to 1080p;
duration: null lets the model pick a length. fal Wan 3 guide.
- H3 Max runs 5 to 15 s at 480p or 768p and renders a 5 s 768p clip in under 3 s wall time. fal H3 Max guide, 2026-09-01; fal H3 Max intro, 2026-08-26.
- H3 Max gives five free 5 s generations a day on its tool page and five more per day in the fal sandbox at up to 15 s. fal H3 vs H3 Max.
- Omni Flash defaults to 8 s and bills $0.03/s at 360p, $0.10 at 720p, $0.15 at 1080p, $0.30 at 4K. Omni llms.txt.
minimax/h3-max/director is a WebRTC realtime stream (fal.realtime.open), not a queue job: 480p/768p, 16:9/9:16/1:1, prompts sent live, memory of 1 to 50 prior segment prompts, $0.02 per generated second promo until Sep 14 then $0.08, 60 s minimum billing per session, sessions over 2 minutes by approval. fal director llms.txt, fetched 2026-09-04. Its example gallery includes a "re-ai-lity dating" stream. fal, fetched 2026-09-04.
- H3 Max's batch endpoints cut inside one generation when the prompt writes "CUT." between timed beats; fal's own example is a three-shot fashion sequence in 10 s. fal, 2026-09-01. H3 Max image-to-video promo pricing ($0.0125 at 480p, $0.02 at 768p) ends September 7 per the i2v page, then $0.05/$0.08. H3 Max i2v llms.txt.
- Kling 3.0 Pro takes a
multi_prompt list with a prompt and duration per shot and returns one stitched clip; $0.112/s no audio, $0.168/s with audio, $0.196/s with voice control. fal Kling 3 llms.txt, fetched 2026-09-04.
- Higgsfield's tested Seedance 2.5 templates run 4 to 24 hard cuts inside a 15 to 30 s generation, and they report a vague LOCATION block as the most common cause of drift between cuts. Higgsfield, fetched 2026-09-04.
- The 30 s montage method: break the script into 2 to 5 s narrative beats, ask for a 15 to 30 s fast-cut montage per beat with extra coverage, generate about 10 variations, cut the best shots in an NLE. beechinour, X, 2026-08-06.
- 30 s beat structure that holds: 0 to 6 s setup, 6 to 14 build, 14 to 24 turn, 24 to 30 end; repeat subject, action, location, camera, style and rules per shot. rewind02, X, 2026-09-03.
- ByteDance: too much content in a time range produces extra cuts or dropped plot, so allocate seconds per shot and keep the timeline gap-free. BytePlus.
- Editorial tradeoff stated by practitioners: single-pass multi-shot gives seam-free continuity and one price per clip, but every cut is a model decision; separate clips give the editor the cut point and let each shot be re-rolled alone. Saga survey of 40 AI filmmakers, 2026; beechinour, X, 2026-08-06.
- Minimum clip length vs 1 to 2 s trailer cuts: generate the full 5 to 10 s, cut the 1 to 2 s that lands; Kroll-style pacing tests show Higgsfield templates with average shot under 1.5 s inside a 30 s pass. Higgsfield, fetched 2026-09-04.
- fal's iteration rule: when 24 s of a 30 s result work, rewrite only the state where the bad action starts and ends; change one thing per re-roll. fal, 2026-08-07.
4. Quality control
- Identity metric used in the reference-video literature: cosine similarity of face embeddings between reference and generated frames ("ID-Sim"). MAGREF, arXiv 2505.23742.
- No fetched practitioner post describes an automated face-embedding gate in a film pipeline; the metric appears in papers and the xsearch synthesis, and fal's own evaluations were human A/B. fal, 2026-08-10.
- InsightFace: standardize on cosine over L2-normalized
normed_embedding, detector plus 5-point alignment to 112x112, buffalo_l (w600k_r50) for server use, antelopev2 for large server; 1:1 thresholds land around 0.30 to 0.45 cosine at FMR 1e-4 to 1e-5 and must be recomputed on your own validation split. InsightFace guide, fetched 2026-09-04.
- AutoMV runs Gemini 2.5 Pro verifier agents on each keyframe candidate (up to 3) and on each clip for script alignment and physical feasibility, with regeneration and fallback; its LLM judge still lags human experts. arXiv 2512.12196, 2025-12.
- OpusClip's production judge: rubric scores on a 0/1/2 scale for hook, content, visual, audio; 250 samples to validate 80% agreement at 95% confidence; 75.2% rubric accuracy on held-out clips; a fixed anchor set rescored 3+ times before any prompt change ships. OpusClip Engineering, fetched 2026-09-04.
- fal hosts
fal-ai/video-understanding (question answering over a video URL) for judge calls without a separate vendor. fal, fetched 2026-09-04.
- fal's own LoRA evaluation was same-seed A/B pairs reviewed by hand, 100 duels, plus blind votes; automated metrics "capture only part". fal, 2026-08-10; fal H3 Max intro, 2026-08-26.
- WaveSpeed's test-log template records prompt version, parameters, reference asset IDs, cost, and failure class (identity drift, camera miss, timing miss, unwanted dialogue, weak audio, artifact). WaveSpeed, 2026-08-26.
- Regeneration rates reported: Seedance 2.5 scenes landing in 1 to 2 generations versus 10+ on prior models. ShamiWeb3, X, 2026-08-06; a 30 s scene on the first try after 8 attempts elsewhere. FutureStacked, X, 2026-08-05; 50+ rejected generations on a Seedance 2.5 prompt. MatthewKadish, X, 2026-09-01; Kling Character ID holds identity in "90%+" of clips, failing on extreme angles, dark light, small faces. aimagicx, 2026; 10 to 20 attempts to pin a character initially. Kling blog, fetched 2026-09-04.
- Across 10,247 tracked generations, iterating one prompt then batching 10 gave 77% keepers versus 26% for blind batches of 100. Cliprise, fetched 2026-09-04.
5. Post
fal-ai/topaz/upscale/video: $0.01 per second up to 720p output, $0.02 to 1080p, $0.08 above 1080p; price doubles at 60 fps; Gaia 2 is half price. fal, fetched 2026-09-04.
- Topaz model families on that endpoint: Proteus (most footage), Artemis (degraded sources), Gaia HQ/CG (rendered), Gaia 2 (animation, 2x), Nyx (denoise), Starlight (diffusion restoration). fal.
fal-ai/seedvr/upscale/video: $0.001 per megapixel of output; 1920x1080 x 121 frames = $0.25. fal.
fal-ai/flashvsr/upscale/video: $0.0005 per megapixel; same clip = $0.125. fal.
- MiniMax H3 (base) offers 2K and 4K as upscales of a 768p render at $0.13 and $0.16 per second; H3 Max stops at 768p. fal H3 vs H3 Max, fetched 2026-09-04.
- Practitioner consensus on X: generate at 480p to 720p, upscale with Topaz Starlight or Astra to 4K, add grain. xsearch synthesis, 2026-09-04, 31 posts.
- Frame interpolation on fal:
fal-ai/film and fal-ai/rife at $0.0013 per compute second; Topaz doubles price for 60 fps output. fal film; fal rife.
- Grade: film grain alone on sharp footage looks fake; apply halation first, a light blur on the base, then grain with its own blur. noamkroll, X, 2026-08-13. Curious Refuge adds a small amount of grain to remove the "AI tinge". Curious Refuge, 2026-04-14. Higgsfield ships a Resolve plugin for AI-assisted LUTs. higgsfield, X, 2026-06-08. Kroll's replies name Filmbox Pro and warn that 1080p YouTube compression eats grain. noamkroll thread.
- Remotion renders parameterized React templates (lower thirds, kinetic type, logo reveals) and ships agent skills and plugins for Claude Code and Codex. Remotion AI docs, fetched 2026-09-04.
- Remotion's license is free for individuals and companies up to three people; larger companies pay. Spinner, 2026-03-20.
- Remotion Lambda splits a render across parallel AWS Lambda functions and stitches the parts. Remotion Lambda docs, fetched 2026-09-04.
- Spinner's shipped alternative: Puppeteer renders HTML/CSS cards to PNG, FFmpeg composites and concatenates by transport-stream copy without re-encoding sources, a final audio-only re-encode removes pops at joins, loudness normalized to -14 LUFS. Spinner, 2026-03-20.
- FFmpeg drawtext is called unreliable for titles in that post; HTML rendering handled fonts, gradients and transparency. same.
- Curious Refuge's edit stage is Premiere Pro or DaVinci Resolve after per-clip Topaz upscaling. Curious Refuge, 2026-04-14.
- Music on fal:
fal-ai/elevenlabs/music $0.60 per output minute rounded up; fal-ai/stable-audio-25/text-to-audio $0.20 per audio; fal-ai/lyria3 $0.04 per audio, fal-ai/lyria3/pro $0.08; fal-ai/minimax-music/v2.6 and fal-ai/ace-step also listed. fal Eleven Music; Stable Audio 2.5; Lyria 3; Lyria 3 Pro; fal explore, fetched 2026-09-04.
- Eleven Music terms: trained on licensed stems, broad commercial use on paid plans, but film, TV and large studio game rights require an Enterprise plan; composition plans allow per-section styles, lyrics and durations from 3 s to 10 minutes. ElevenLabs Music API page, fetched 2026-09-04.
- Suno: free plan songs are owned by Suno and non-commercial; Pro ($8/mo billed yearly, 2,500 credits) and Premier ($24/mo, 10,000 credits) grant ownership and a commercial license for new songs; v5.5 is the current model on the pricing page. Suno pricing, fetched 2026-09-04; Suno help, edited ~2026-01. No official Suno API page was fetched; a third-party listing claims one. APIRank, fetched 2026-09-04.
- Udio: terms forbid downloading or distributing Output on any platform, and forbid commercial exploitation, so Udio cannot supply a trailer track. Udio ToS, fetched 2026-09-04; pricing $8 to $24/mo. Udio pricing.
- Lyria: DeepMind's page says tracks up to 3 minutes, via Flow Music and Gemini. DeepMind, fetched 2026-09-04.
- SFX:
fal-ai/elevenlabs/sound-effects/v2 $0.002 per second; ElevenLabs direct $0.12 per minute, royalty-free. fal; ElevenLabs API pricing.
6. Prompting conventions for realistic humans
- Structure that all three fal guides converge on: FORMAT line (duration, ratio, single take or cuts), REFERENCE ROLES, STARTING STATE, TIMELINE with second ranges, CAMERA with screen position and trigger event, CONTINUITY invariants, AUDIO, ENDING STATE, CONSTRAINTS. fal Seedance guide, 2026-08-07; Higgsfield GLOBAL STYLE/SCENE/CHARACTERS/LOCATION/FIRST FRAME AND BLOCKING/OPTICS/CAMERA/AUDIO; Dreamina 8-part formula, fetched 2026-09-04.
- Camera language ByteDance says to write literally: shot size (wide, medium, MCU, close-up), movement (push in, pull out, pan, track, follow, orbit, tilt, handheld shake), angle (low, overhead, first-person), techniques (long take, dolly zoom, speed ramp). BytePlus. One camera action per beat. WaveSpeed, 2026-08-26; fal Wan 3 guide.
- "Dynamic" is useless; put the subject in a named third of frame and tie the pan to a visible event. fal Seedance guide, 2026-08-07. Comma lists of aesthetic words do not help Kling 3; 2 to 5 sentences per shot is the sweet spot. fal Kling 3 guide.
- Confessional setup phrasing that a reality-TV prompt library uses: "a person in a styled interview chair speaking to camera under soft key light, blurred decorated set behind, medium shot, shallow focus"; drama beats: "handheld camera snapping between faces". Morphic, fetched 2026-09-04. fal's H3 Max courtroom example adds the delivery note ("quiet and unhurried") and "looks straight down the lens". fal, 2026-09-01.
- Anti-polish block used by Higgsfield and ByteDance examples: "shot on Arri Alexa Mini LF, 35 mm cinema lens, film grain, authentic skin texture, natural lifelike performance, subtle micro-expressions, real adult facial bone structure, no excessive beautification or skin smoothing". BytePlus; "Skin unretouched, pores, chapped lips, no beauty retouch, no digital smoothing, no CGI sheen". Higgsfield. For a phone-camera look: "autofocus hunting once, exposure shifting, handheld wobble, no colour grade". fal H3 vs H3 Max, fetched 2026-09-04.
- Cause before reaction, contacts named, hidden objects re-described on re-entry, fluid or hair motion given an end state. fal Seedance guide, 2026-08-07.
- Extra people: lock the head count ("exactly four, never five, no duplicate member") in a POSITIVE LOCKS block. Higgsfield; "no duplicated character, no morphing" in Wan's fal example. fal Wan 3.0 Prime page.
- Known failure modes: on-screen text (ByteDance flags text accuracy as still improving; Higgsfield "no readable text"), hands (practitioner negatives list extra or fused fingers), teeth and eye contact (practitioner prompt adds "eyes looking directly into the camera" on the reference sheet and "natural eye contact" in video). WaveSpeed citing Alibaba release notes, 2026-08-26; meAsifAi, X, 2026-08-27.
- Prompt length: satisfaction peaked at 21 to 50 words and fell above 100 in one platform's data, and a failing 300-word prompt should be cut to 90 before adding detail. Cliprise; WaveSpeed, 2026-08-26.
- Safety controls exposed on fal: Wan 3.0 Prime and H3 Max have
enable_safety_checker (default true); Veo 3.1 has safety_tolerance 1 to 6 (default 4); Seedance 2.5 endpoints expose no safety parameter. Wan llms.txt; H3 Max llms.txt; Veo llms.txt; Seedance llms.txt.
- Romance and drinking in fetched, rendered examples: Higgsfield's tested Seedance 2.5 drama template has two lovers in blankets saying "I will always love you" with a STILLNESS LOCK ("nobody hugs, nobody kisses, nobody's hands leave the blankets") and one scripted tear at 13.2 s. Higgsfield, fetched 2026-09-04.
- fal's Wan 3 dialogue example is a two-hander over a sheaf of papers written as STYLE, CHARACTERS, LOCATION, TIMELINE and SOUND blocks with lines in quotes. fal Wan 3 guide.
- The AI dating soap's viral episode (girlfriend flashes her boyfriend as police take him away) rendered on Seedance 2.5 and reached 4.2M views, so mild adult comedy passes when written as staged action. GuntherWrite, X, 2026-08-29.
- These examples specify behavior, timing and camera and avoid anatomical or explicit words; none of the fetched pages lists Seedance 2.5's refusal triggers (see Not verified).
7. Recommended pipeline for Manufactured Love
- Cast bible: 8 characters, one text block each in fixed order (age, build, face, hair, skin, wardrobe with hex colors, one identifying prop), plus a "never change" list. Elser; Kling blog.
- Reference sheets: Nano Banana Pro or Seedream 5 for a front portrait, then edit endpoint for profile, 3/4, back, and a 3x2 sheet with three large head crops; separate wardrobe plate on white; confessional-chair location plate and villa plates. Reconstruction test before sign-off. Expect 10 to 20 image generations per character. meAsifAi; fal Wan 3 guide; Kling blog.
- Voice bible: one ElevenLabs Voice Design voice per character, saved to a slot; write dialogue as v3 with audio tags; export one WAV per line. Keep a table of character, voice ID, design prompt, reference WAV. ElevenLabs Voice Design docs; JoshDaws.
- Shot list: 45 to 55 shots, each with model, duration (5 s on H3 Max, 5 s on Wan Prime, 8 s on Omni, 6 to 10 s on Seedance), references by role, timeline, camera, continuity, audio, ending state; every cut ends in the pose the next begins. maverickecom; fal Seedance guide.
- Draft pass at 480p/768p: Wan 3.0 Prime reference-to-video for two-person and group villa shots (10 image refs), H3 Max image-to-video at 768p for single-character b-roll from keyframes, Seedance 2.5 reference-to-video at 480p only for shots needing more than 10 refs or over 15 s. Three takes per shot. Log seed, prompt version, refs. fal Seedance page; WaveSpeed.
- Confessionals: generate the confessional-chair still per character once; drive it with the ElevenLabs line through
fal-ai/bytedance/omnihuman/v1.5 at 720p, or as native Seedance 2.5 with @Audio1 as the voice reference when a camera move is wanted. Two takes per line. fal OmniHuman guide; fal Seedance page.
- QC gate: sample every 6th frame, InsightFace
buffalo_l cosine against each character's front crop, threshold set from a validation set of accepted takes (start around 0.40), fail if any sampled frame with a detected face drops below the floor or no face is found in a talking shot; then a Gemini or fal-ai/video-understanding rubric (0/1/2) for hands, text, extra people, eye line, lip motion; then a contact sheet for human pick. InsightFace guide; OpusClip; AutoMV.
- Re-roll only failing shots, one change per re-roll, same seed when only resolution changes. fal Seedance guide.
- Final pass: rerun accepted shots at 1080p on Wan 3.0 Prime and Omni Flash; H3 Max shots stay 768p and go through
fal-ai/topaz/upscale/video to 1080p ($0.02/s). fal Topaz.
- Post: DaVinci Resolve for the cut; halation, light blur, 35 mm grain; Remotion for lower thirds and title cards from one parameterized template; FFmpeg stream-copy concat and a final audio pass at -14 LUFS. noamkroll; Remotion; Spinner.
- Music and SFX: Stable Audio 2.5 or Lyria 3 Pro on fal for the trailer bed at draft; for release, confirm rights (Eleven Music needs Enterprise for film/TV; Suno needs Pro or Premier). ElevenLabs SFX for stings and whooshes. ElevenLabs Music API page; Suno help.
Expected regen: 3 takes per shot at draft (keep 1 of 3), 1.5 takes per accepted shot at final; talking shots 2 takes per line. Basis: 1 to 2 gens per scene reported on Seedance 2.5, 10+ on older models, 77% keeper rate with iterate-then-batch, 90%+ identity hold with good references. ShamiWeb3; Cliprise; aimagicx.
Per-shot cost, 5 s clip unless noted, fal list prices fetched 2026-09-04:
| Model |
Draft (480p or 768p) |
Final (1080p) |
Source |
| Wan 3.0 Prime r2v |
$0.34 at 480p, $0.70 at 720p |
$1.40 |
fal |
| H3 Max i2v |
$0.06 at 480p / $0.10 at 768p promo to Sep 7; $0.25 / $0.40 list |
768p + Topaz $0.10 = $0.50 list |
fal; Topaz |
| Seedance 2.5 r2v / i2v |
$1.10 at 480p, $2.37 at 720p |
$5.82 (i2v/t2v only) |
fal |
| Gemini Omni Flash 1.1, 8 s |
$0.24 at 360p, $0.80 at 720p |
$1.20 |
fal |
| Veo 3.1 r2v, 8 s with audio |
$3.20 at 720p |
$3.20 |
fal |
| OmniHuman 1.5 talking head |
$0.80 at 720p |
$0.80 at 1080p |
fal |
| Kling lipsync on existing clip |
$0.014 |
$0.014 |
fal |
| sync v3 on existing clip |
$0.67 |
$0.67 |
fal |
| SeedVR2 upscale to 1080p |
n/a |
$0.25 |
fal |
Trailer totals at the regen rates above, 50 shots: draft 150 takes on Wan Prime 480p about $51 (H3 Max 768p list about $60, Seedance 480p about $165); final 75 takes on Wan Prime 1080p about $105 (Omni 1080p about $90, Seedance 1080p about $437); 20 talking lines x 2 takes on OmniHuman about $32; ElevenLabs v3 for about 1,500 characters x 10 revisions about $1.50; Topaz on a 120 s cut $2.40; music $0.08 to $1.20 per candidate. Reference images are not priced here.
Not verified
- Seedance 2.5's actual moderation trigger terms. The only cited post on refusals reports repeated rejections without naming the trigger; two xsearch answers on this topic returned no citations and were discarded. MatthewKadish, X, 2026-09-01.
- Whether the H3 Max promo ends September 7 (image-to-video page) or September 14 (director page); both pages were fetched 2026-09-04 and disagree. i2v; director.
- Seedance 2.5 1080p on the reference endpoint: the fal model page lists 480p and 720p only, while the i2v/t2v pages and Higgsfield say 1080p. fal r2v page.
- The 8 to 18% / 25 to 55% regen bands and the 0.50 to 0.65 cosine bands that appeared in one uncited xsearch answer; not used above.
- Typical regen counts on Wan 3.0 Prime, H3 Max and Omni Flash; only Seedance 2.5 has cited hit-rate reports.
- Gemini Omni Flash 1.1 reference image count on fal; the llms.txt does not state a cap.
- Any public per-character LoRA case study on H3 or Wan 3 for a multi-character short; none found.
- Kling 3.0
multi_prompt behavior with 6+ characters; fal's guide claims 3+ without numbers.
- Suno's official API and its terms; only a third-party listing was fetched.
- Reddit threads (r/Seedance_AI drift, face censorship, Seedance 2.5 practical guide) could not be fetched in any form and are absent.
- The cost table uses list prices without reference-video billing (Seedance x0.6 with video input; H3 Max token billing past 4,096 tokens) and without Nano Banana Pro image costs.