- family
- character-consistency
- modality
- text/image -> image -> identity-lock -> video+voice -> polish
- whyMultiModel
- Chains at least three distinct generative systems: an image generator (Soul or Nano Banana) for the base character still, Soul ID as a trained identity-lock layer that persists the face across generations without re-uploading a reference each time, and Veo 3 as a separate video-and-voice model that produces native lip-synced dialogue directly (no separate TTS or lip-sync model bolted on afterward). This differs from reference-locked-character-consistency (InstantID/IP-Adapter face lock on still images only, no video stage) and from voice-clone-to-talking-head / tts-video-diffusion-lipsync-chain (both of which pair a separate zero-shot TTS clone with a separate lip-sync/talking-head model rather than a video model with native audio generation).
- steps
- [object Object], [object Object], [object Object] +1 more
- controls
- Soul ID reference photo set for identity training; text prompt describing motion, setting and dialogue for Veo 3; VFX and camera-motion node parameters; export aspect ratio (9:16 or 16:9)
- exampleStack
- Soul or Nano Banana (still) -> Soul ID (identity lock) -> Veo 3 (talking video + voice) -> Higgsfield VFX/Camera Motion (polish)
- useCases
[recurring branded spokesperson videos across campaign variations][UGC-style talking avatar ads without re-shooting per script][consistent narrator/host across a multi-episode video series]
- pitfalls
- Requires a Higgsfield subscription tier with Veo 3 access; identity lock quality depends on the base reference image; native Veo 3 audio/dialogue means separate TTS or lip-sync tooling is not used, so voice style is constrained to what Veo 3 can render from the prompt rather than a cloned voice.