- family
- audio-driven
- modality
- image -> text (script) -> audio (cloned voice) -> video
- whyMultiModel
- No single model both writes a grounded performance script from an image, clones/synthesizes a matching voice, and renders a lip-synced video; the chain needs an image-conditioned script generator, a separate voice-cloning TTS model, and a distinct audio-conditioned video generator.
- steps
- [object Object], [object Object], [object Object]
- controls
- source image, script/expression tags from the LLM stage, voice selection or cloned voice sample, scene description prompt, LTX-2.3 duration and lipsync strength
- exampleStack
- Gemini (script) -> ElevenLabs Instant Voice Clone (narration) -> LTX-2.3 (video + native lipsync)
- useCases
[product explainer videos from a single photo][testimonial-style UGC ads][social media talking posts without filming]
- pitfalls
- LTX-2.3's native lipsync can drift on longer scripts or fast speech; cloned voice quality depends on the reference sample length and clarity, and mismatched scene description vs. script tone produces awkward framing