A one-stage model animates a full human (not just the face) from audio, scaled up by conditioning on portrait, pose, and body signals together so motion and lip-sync emerge from a single network.
Properties
family
audio-driven
modality
image/video + audio -> human-animation video
whyMultiModel
Full-body co-speech motion, gestures, and lip-sync run on different timescales; OmniHuman unifies them but still pairs with speech synthesis and upscaling in a real pipeline.
steps
[object Object], [object Object], [object Object]
controls
Source portrait or video; audio; body and pose conditioning; motion intensity.
exampleStack
TTS voice clone -> OmniHuman -> 2x video upscale.
useCases
[Presenter and spokesperson videos][Multilingual dubbing with motion][Social-content avatars]
pitfalls
Hands and fine gestures can smear under fast motion; low-resolution inputs upscale with artifacts.