Loading...
A source video is passed through a preprocessor that extracts a structural control signal (depth map, Canny edges, or OpenPose skeleton). That signal is fed as conditioning to a video diffusion model which generates new video that preserves the original motion and geometry while applying a new appearance, style, or character. The result is a restyled video with motion inherited from the reference footage.