R6. Image prompting for photorealistic reality-TV stills, as of 2026-09-04
Source of truth: four parallel research passes (firecrawl web search/fetch + xsearch) run 2026-09-04 against fal.ai model docs, lab prompting guides, Artificial Analysis / Arena leaderboards, and practitioner threads. Current pipeline is fal-ai/bytedance/seedream/v5/lite t2i + edit; client reports outputs read too clean, too symmetrical, plastic skin, render-like backgrounds, generic faces.
1. Per-model prompting guides
ByteDance Seedream 5 Lite (fal-ai/bytedance/seedream/v5/lite, /edit)
- Natural language beats keyword stacking; tokens like "masterpiece, best quality, 8K" distract the model's chain-of-thought reasoning pass before pixel generation. fal Learn, 2026-03-19
- Structure: subject+action, then setting, then style/mood as adjectives on the whole scene ("a moody noir photograph of a rainy street," not a tag list); 2 to 4 sentences is the sweet spot, content past ~50 words gets deprioritized. Same source.
- Edit endpoint takes up to 10 reference images, addressed positionally in prose as "Figure 1," "Figure 2." Composition "works best when reference images have similar lighting conditions"; mismatched refs are handled but look less natural. Same source.
- No seed or negative-prompt control exposed on either endpoint. Same source.
ByteDance Seedream 5 Pro (fal slug reported as bytedance/seedream/v5/pro — confirm against the live catalog before wiring)
- "Reproduces real-world lighting, materials, and skin textures, balancing CG expressiveness with photographic quality"; facial lines and rough skin read three-dimensional, matte lighting transitions softly. ByteDance Seed blog, undated ~mid-2026, fal.ai/seedream-5.0, fetched 2026-09-04
- Edit endpoint: up to 10 refs, region-precise sketch/point/lasso edit, layer separation. Multi-person example assigns roles by number — "Combine the people from Images 2 to 6 into a group photo referencing the positioning in Image 1" — numbered-image role assignment works for identity+composition, not just style. Same sources.
Seedream 4.5 (predecessor, still documents core technique)
- "Order matters": lead subject, then style, then composition, then lighting, then technical params. fal Learn, undated, fetched 2026-09-04
- Concrete film-language example: "shot on expired Kodak Portra 800, pushed two stops... visible grain, halation around light sources, slightly lifted blacks." Replicate blog, undated, fetched 2026-09-04
- "Example-based editing": give a before/after pair (Image 1 to Image 2) plus a new target (Image 3), prompt "reference the change from Image 1 to Image 2, apply the same operation to Image 3" — transfers style without describing it in words. Same source.
Google Nano Banana Pro / Nano Banana 2 (Gemini 3 Pro Image / Gemini 3.1 Flash Image)
- Accepts up to 14 reference images in one prompt; Gemini 3 Pro Image splits the budget by role — up to 6 high-fidelity object images, up to 5 character-consistency images, up to 3 style-reference images. Gemini 3.1 Flash Image: up to 10 objects + 4 characters, no dedicated style slot. Google Cloud Blog, 2026-03-05, ai.google.dev, fetched 2026-09-04
- Generation formula: [reference images] + [relationship instruction] + [new scenario]. "Prompt like a creative director": name lighting explicitly ("three-point softbox," "harsh chiaroscuro," "golden hour backlighting"), name camera/hardware explicitly ("GoPro" for distorted immersive, "cheap disposable camera" for raw nostalgic flash), name film stock/grade explicitly. Same sources.
- All outputs carry C2PA content credentials + SynthID watermark — not visually detectable but present in file metadata; relevant if refs must pass provenance checks. Same source.
Black Forest Labs FLUX 2 Pro
- Loose template:
[SUBJECT], [LOCATION], [STYLE], [CAMERA SETTINGS], [LIGHTING], [COLORS], [EFFECT], [ADDITIONAL ELEMENTS]; natural language preferred, refine one change at a time. BFL docs, undated, fetched 2026-09-04
- Direct quote on photorealism: "prompt the model as if a real photo is being captured in the moment. Use photography language (lens, lighting, framing) and explicitly ask for real texture (pores, wrinkles, fabric wear, imperfections)." Same source, usecases_t2i_photorealistic
- Edit endpoint: up to 8 reference images via API, 10 in the Playground; numbered role assignment confirmed by BFL's own example — "Use Image 2 as the location. Insert only the ice skates from Image 1 into Image 2, with the decorations and evening lighting vibe from Image 3." Same source, prompting_editing_overview
- Style transfer pattern: preserve-list + transform-instruction, e.g. "preserving the subject's skin-tone values... exact pose, proportions, lighting, and composition." Same source.
OpenAI gpt-image-1.5
- OpenAI's current flagship is now
gpt-image-2; the live prompting guide centers on it and lists gpt-image-1.5 as the prior model still selectable. OpenAI cookbook, undated, fetched 2026-09-04 — flag to team, the brief specified 1.5.
- Photorealism guidance is near-identical wording to BFL's: "prompt the model as if a real photo is being captured in the moment... explicitly ask for real texture... avoid words that imply studio polish or staging." Same source.
- Include the literal word "photorealistic" to engage photorealistic mode; detailed camera specs "may be interpreted loosely — use them mainly for high-level look... rather than exact physical simulation," a looser coupling than FLUX. Same source.
- Multi-image edit: reference by index + description ("Image 1: ...; Image 2: ...") and state the interaction; no per-image role parameter. Same source.
- Anti-cinematic example in the guide itself: "avoid cinematic lighting, dramatic color grading, or stylized composition... grounded, authentic, and unstyled, as if captured in a real moment." Same source.
2026 leaderboards
- Artificial Analysis text-to-image Elo, fetched 2026-09-04: 1. GPT Image 2 (high), 1371; 2. MAI-Image-2.6 (Microsoft), 1345; 3. Reve 2.1, 1324; 4. Nano Banana 2, 1320; 6. gpt-image-1.5, 1304; 9. Nano Banana Pro, 1296; 13. Seedream 5.0 Pro, 1279; 20. FLUX.2 [max], 1224; 33. FLUX.2 [pro], 1207; 43. Seedream 5.0 Lite, 1196. artificialanalysis.ai/image/leaderboard/text-to-image
- Arena leaderboard corroborates the ordering. arena.ai/leaderboard/text-to-image, fetched 2026-09-04
- GPT Image 2 beats all four models named in the brief on general Elo; neither board isolates a photorealism-specific or a people/candid-specific sub-metric, so this is suggestive for this use case, not dispositive — worth an eval pass against the client's reference set before switching pipelines.
2. Realism techniques verified in 2026
- Camera+lens+aperture naming is still the highest-leverage single lever, but it functions as a style signal, not literal optics — brand names bias color/contrast toward what that brand looks like in training data. roo.beehiiv.com, 2026-05-21, Medium, 2026-04-15
- Generic quality words ("ultra-realistic," "8K," "cinematic," "masterpiece") are dead weight and can trip an "aesthetic mode" that produces exactly the glossy look the client is complaining about; describe the capture, not the beauty. Envato, 2026-08-13, Hedra, fetched 2026-09-04
- Studio-grade gear names (Arri Alexa, RED, Leica) read as "trying too hard"; humble gear (phone, Ricoh GR III, old Canon AE-1) reads more real. Same Hedra source.
- "Amateur/unremarkable iPhone photo" framing produced markedly more realistic output than "professional photography" framing in a direct practitioner comparison. Reddit r/singularity, fetched 2026-09-04
- Real phone-photo signature: cheap on-camera flash, a hotspot on the face, near-black background falloff, slight motion blur and tilt — "a photo reads as real when it was taken badly." Same Hedra source.
- Golden hour is now an AI signature itself — every model defaults to warm raking light, so it reads as a tell; flat/overcast or plain unflattering light reads as more real. Same Hedra source.
- Film stock/grain still helps generally, but Portra 400 and Cinestill 800T are now over-prompted enough to read as "trying to look like film"; Hedra recommends Kodak Gold 200, Ektachrome, or Ilford HP5 instead — this contradicts Envato's guide, which still recommends Portra 400 by name, an active disagreement with no adjudication found. [Hedra, fetched 2026-09-04]; [Envato, 2026-08-13]
- ISO-linked grain language ("shot at ISO 1600, natural film grain") is standard across sources. Miraflow, 2026-04-30
- Imperfect framing (tilted horizon, off-center, motion blur) is consistently recommended but has no controlled test behind it — practitioner consensus, not measured. [Envato, 2026-08-13]; [Hedra, fetched 2026-09-04]
- A single, named, plausible light source (window light, noon sun, on-camera flash) reads more real than "cinematic lighting," which averages into contradiction. Reddit r/generativeAI, fetched 2026-09-04, [Envato, 2026-08-13]
- Skin: "visible pores, subtle asymmetry, subsurface scattering" beats "photorealistic skin"; avoid "flawless," "polished," "beauty skin" — these specifically trigger the plastic look. [roo.beehiiv, 2026-05-21]; Instagram, fetched 2026-09-04, undated post
- Positive framing beats negation generally — describe concrete age/occupation/skin detail rather than negating ("not a model"); models reportedly under-weight negated identity descriptions. [Hedra, fetched 2026-09-04]
- AI-symmetry fix: web sources consistently say prompt "subtle facial asymmetry" to break the generic-face look. [roo.beehiiv, 2026-05-21]; Reddit r/StableDiffusion, fetched 2026-09-04 — an xsearch pass returned the opposite claim (prompt symmetry to fix deformed faces) with zero citations; that answers a different failure mode (anatomical error, not generic-face look) and is uncited, so weight the web sources.
- "Not a render" / negative meta-instructions: used in production marketing copy for AI photo tools, but no source directly tests it against a version without it. Could not find evidence either way on backfire risk; flag as unresolved rather than guessed. getmatches.ai, fetched 2026-09-04
- Prompt length/ordering is the weakest-evidenced sub-topic here. One source asserts camera-language placement near the start or end of the prompt matters; no dated, sourced comparison of short vs. long prompts per specific model was found. [roo.beehiiv, 2026-05-21]
- Seed/resolution effect on perceived realism: no credible 2026-specific source found; do not assert a claim.
- Real production clutter sells the documentary look: workbench clutter, harsh shop fluorescents mixed with soft daylight, off-center framing, subject looking at hands not camera; drop any detail you won't art-direct, since ungoverned details default to generic. [Hedra, fetched 2026-09-04]; crumbs, drying dishes, coffee-ring stains, dust in a light beam as lived-in cues. [Miraflow, 2026-04-30]
3. Style-reference workflows: real photo as style ref, character sheet as identity ref
- Seedream 5 (Lite and Pro) edit: up to 10 refs, addressed positionally in prose ("Figure 1," "Figure 2"), no JSON role field; composition and style transfer are both prompt-driven, similar lighting across refs recommended. [fal Learn, 2026-03-19]; blog.fal.ai, undated post published after 2026-02
- Nano Banana Pro is the only model of the four with a named, separate style-reference budget (up to 3 images) distinct from character/identity budget (up to 5 images) — the strongest lever available for "match this real press photo's style, keep this identity." [ai.google.dev, fetched 2026-09-04]
- Documented Nano Banana role-splitting syntax: "Take the [element from image 1] and place it with the [element from image 2]... with the lighting and shadows adjusted to match." Legacy
gemini-2.5-flash-image works best at ≤3 input images; gemini-3-pro-image supports 5 at high fidelity, 14 total mixed. Same source.
- FLUX 2 Pro edit: up to 8 refs via API (10 Playground); numbered role assignment confirmed in BFL's own example — one image as location, one as inserted object, one as lighting-vibe donor, three distinct roles in a single prompt. [docs.bfl.ml, fetched 2026-09-04]
- gpt-image-1.5/2 edit: official docs show multi-image compositing by index+description with no documented identity-vs-style role split. A widely-read practitioner guide documents a two-image pattern instead: first image locks identity ("preserve my facial identity, proportions, age, and expression fidelity exactly, do not stylise my face"), second image sets style/composition. [OpenAI cookbook, fetched 2026-09-04]; Substack, 2025-12-21 — practitioner pattern, not OpenAI-published.
- Cross-model: role assignment goes through prose ordinals everywhere except Nano Banana Pro, which has typed input slots. No source gives a quality-degradation curve for reference count — none of the four docs states where added refs stop helping.
- An xsearch pass on this exact sub-topic returned fabricated, uncited output (invented
role/ref_weights JSON fields, conflated "Banana.dev" GPU cloud with Google's Nano Banana) with zero sources used; excluded from this doc, flagged here as an xsearch reliability gap on this specific query.
4. Reality-TV-specific look: shoot it like the real thing, not "cinematic"
- Love Island UK's on-set photographer (2018 finale) shot two Nikon D5 bodies, 70-200mm f/2.8 and 200-500mm f/5.6 for the live final, 50mm f/1.4 and 24-70mm f/2.8 for couple portraits, Nikon SB-5000 flashguns on wireless trigger, images FTP'd live to a London editor same-night. Amateur Photographer, 2018-08
- Pre-villa cast entrance photos are shot in under 10 minutes in an improvised location (a shipping container, per one 2021 islander's account), directed poses, no retakes or preview. OK! Magazine, 2023-06-21
- Fans and cast repeatedly flag harsh, unwanted on-camera flash and unflattering angles on official promo shots versus how cast look on the show itself — a real, visible artifact of the flash-lit quick-turn workflow, not a look production hides. The Tab, archived 2022-07
- Too Hot to Handle's unit stills photographer is credited as Ana Blumenkron; no gear/lighting interview found for the show specifically, closest verified analog is the Love Island workflow above (both sun/villa dating formats). IMDb, fetched 2026-09-04, no page date
- The Bachelor/Bachelorette: no photographer interview or gear source found in this pass. Do not extrapolate a specific look for this franchise; treat as an open gap.
- General unscripted-TV stills practice: mirrorless/DSLR bodies (Canon R5 cited), 50mm/85mm primes for character portraits, 24-70mm/70-200mm zooms for coverage; quick pull-aside portraits during production borrow a spare gaffer light rather than a dedicated rig; full lighting setups are reserved for separate seamless-backdrop "gallery day" shoots. rekhagarton.com, undated, fetched 2026-09-04
- On scripted sets, unit stills photographers avoid flash while cameras roll (silent mirrorless, no flash, to not disrupt sound/picture) — the opposite convention from dating-show publicity portraits, which are shot off to the side and do use flash. The Phoblographer, 2021-11-06
- Practical implication: "candid villa still" (available light, no flash, imperfect, shot during filming) and "cast promo portrait" (flash-lit, directed, slightly harsh) are two different photographic registers — do not use one prompt template for both.
- Anti-pattern to avoid: a commercial AI-photo product sells "Love Island bombshell" presets described as glossy, golden-hour, wet-hair, sun-lounger, "photorealistic fashion magazine cover" — this is the over-produced cliché that likely explains why the client's current output reads fake; real promo shots are flash-harsh and often unflattering, not glossy. myaiphotoshoot.com, fetched 2026-09-04, undated
5. Prompt templates
All four assume the style-reference workflow from §3: a real press still as the style/lighting/grade reference, a character sheet as the identity reference, roles stated explicitly per each model's syntax above.
(a) Cast promo portrait — flash-lit, directed, on-location
- Seedream 5 Pro/Lite: "Figure 1 is a lighting and color-grade reference: on-camera flash, hard fill, slightly blown highlights, tropical dusk background. Figure 2 is [name], keep facial identity exact. Match Figure 1's flash-lit look on Figure 2, three-quarter pose, poolside backdrop, slight motion blur on hair from flash sync." Avoid: "cinematic," "editorial," "8K," "flawless skin."
- Nano Banana Pro: style reference = Image 1 (real press still, flash-lit grade); character reference = Image 2 ([name]). "Apply the lighting and color grade of the style reference to the character reference: on-camera flash fill against hard sun, saturated but real color, slight highlight clipping on skin, three-quarter pose, villa poolside." Avoid: "beauty lighting," "soft glow," symmetrical framing.
- FLUX 2 Pro: "Use Image 2 as the subject and identity source, preserve exact face. Use Image 1 as the lighting/color-grade reference: harsh on-camera flash, hard shadow behind subject, slightly overexposed highlights, compact point-and-shoot look, ISO 800, flash sync." Avoid: "Arri Alexa," "cinematic grade," "professional photography."
- gpt-image-1.5: "Image 1: [name], preserve facial identity, proportions, expression exactly. Image 2: real press-photo style reference. Apply Image 2's on-camera-flash lighting and color grade to Image 1's subject: hard flash fill, blown highlight edge, poolside dusk background, no retouching, no glamorization." Avoid: "studio polish," "staged," "cinematic."
(b) Candid villa still — available light, unposed, mid-action
- Seedream 5 Pro/Lite: "Documentary photograph, available light only, midday sun through villa windows, subject caught mid-conversation not looking at camera, slightly off-center framing, visible pores and uneven skin tone, real clutter in background — towels, drink cups, sunscreen bottle." Avoid: "golden hour," "portrait mode," "bokeh," symmetry.
- Nano Banana Pro: identity ref = Image 2. "Candid moment, natural window and midday light only, unposed, subject mid-laugh looking away from camera, imperfect framing, background villa clutter (charging cables, poolside towels), match the harsh flat midday light of a real phone snapshot." Avoid: "warm backlighting," "cinematic," "flawless."
- FLUX 2 Pro: "Real photo language: iPhone snapshot, flat midday light, slight motion blur, off-center crop cutting part of the frame, subject not posed, visible skin texture, sun glare on lens edge." Avoid: "35mm f/1.4 bokeh," "professional," "masterpiece."
- gpt-image-1.5: "Photorealistic candid photograph, taken on a real camera mid-conversation, unposed, looking away from camera, harsh midday sun, no color grading, no glamorization, honest and unstyled, real villa clutter in background." Avoid: "cinematic lighting," "dramatic color grading," "stylized composition" — a direct anti-pattern named in OpenAI's own guide.
(c) Crew documentary still — behind-the-scenes, working set
- Seedream 5 Pro/Lite: "Behind-the-scenes documentary photograph, crew member adjusting a boom mic near a contestant, shot from the side, harsh on-location work light mixed with daylight, visible cables and stands in frame, grainy, ISO 1600 look, unposed candid gesture." Avoid: "cinematic," "hero shot," clean symmetric composition.
- Nano Banana Pro: style reference = real BTS press photo (mixed practical + daylight, workmanlike framing); subject = crew member mid-task, not looking at camera, gear and cable clutter visible, slightly tilted horizon. Avoid: "golden hour," "beauty lighting."
- FLUX 2 Pro: "Documentary-style photograph, real texture (fabric wear, sweat, dust), mixed color temperature light (tungsten work light and daylight), off-center, crew equipment clutter in frame, imperfect framing." Avoid: "Leica," "shallow bokeh portrait," "polished."
- gpt-image-1.5: "Photorealistic, unstaged behind-the-scenes photograph, crew working, mixed practical lighting, real texture and imperfections, no glamorization, grounded and authentic as if caught in a real moment." Avoid: "cinematic," "dramatic lighting," "stylized."
(d) Location plate — villa exterior, pool deck, resort grounds
- Seedream 5 Pro/Lite: "Wide establishing photograph of villa pool deck, midday hard sun, real slightly blown-out sky, visible pool furniture and towels left out, handheld framing, not perfectly level horizon, no people." Avoid: "drone cinematic," "golden hour," "HDR."
- Nano Banana Pro: "Location reference photograph, hard tropical midday sun, real color (saturated but not stylized), pool deck with everyday clutter (loungers, drinks, towels), slightly tilted handheld framing." Avoid: "aerial cinematic shot," "dramatic sky," "symmetric composition."
- FLUX 2 Pro: "Wide-angle location photograph, harsh midday light, real environmental wear (water stains, faded furniture), off-center framing, phone-camera lens distortion at edges." Avoid: "RED camera," "cinematic wide," "perfect symmetry."
- gpt-image-1.5: "Photorealistic establishing photograph of a resort pool deck at midday, real unstyled light, everyday clutter, handheld framing, no drone shot, no dramatic sky, no color grading." Avoid: "cinematic," "aerial," "stylized composition."
Not verified
Exact fal.ai model slug for "Seedream 5 Pro" (used bytedance/seedream/v5/pro above; R4 only confirmed v5/lite and v4.5/* as active in the live catalog — check before wiring). Whether GPT Image 2 or MAI-Image-2.6 beat the current pipeline specifically on people/candid photorealism, as opposed to general Elo — no sub-benchmark exists. Publish dates for several fal.ai and docs.bfl.ml pages (reported as "fetched 2026-09-04" where the page carries no visible date). FLUX 2 Pro / Kontext community-thread depth beyond BFL's own docs (one Reddit fetch was blocked; BFL's docs covered the gap used here). Whether "not a render" as a literal prompt phrase helps, hurts, or is a no-op on any of the four models — no direct test found either way. Seed and resolution effects on perceived realism — no 2026-specific source found. Location-plate-specific shooting conventions for any of the three named shows — no source found; section 4's plate guidance is inferred from general location-still practice, not verified per show. Bachelor/Bachelorette promo photography gear, lighting, or photographer credits — no source found in this pass. The Clifton Prescod / Love Island: Beyond the Villa key-art crew claim was sourced via xsearch synthesis of an Instagram post, not independently opened — treat as plausible only.