Loading...
A text-to-video (or image-to-video / pose-to-video) model first generates the avatar's body motion and scene, and a separate real-time latent-space lip-sync model then re-renders just the mouth region at 30+ FPS to match streaming audio, decoupling body/scene generation from low-latency mouth synchronization.