- family
- multi-shot
- modality
- text concept -> storyboard image grid -> multi-shot video
- whyMultiModel
- Image models hold composition, lighting, and character silhouette across many panels in one pass; video models are better at motion and camera dynamics but drift on composition if asked to invent both from text alone. Splitting what it looks like from how it moves across two model families is what stabilizes multi-shot output.
- steps
- [object Object], [object Object], [object Object]
- controls
- storyboard panel order/count, per-shot text description, reference-image weighting (omni reference), timestamped multi-shot prompts, anchor-frame selection
- exampleStack
- Nano Banana 2 (storyboard grid + detail fix) -> Kling 3.0 (multi-shot, omni reference)
- useCases
[Ad and trailer previz][Narrative shorts with consistent characters across scenes][Product commercials with multiple staged shots]
- pitfalls
- Consistency comes from a strong anchor frame, not the grid itself; a weak or ambiguous panel drifts its downstream shot, and product/face detail lost during grid rendering silently carries into the animated shot unless fixed first.