Loading...
A video of a person and a flat garment reference image are fed into a diffusion-transformer model that disentangles garment texture/print from pose and injects it frame-by-frame, producing a video of the person wearing the new garment with consistent fabric detail and motion.
Source: https://arxiv.org/abs/2505.21325