Loading...
A subject is segmented in one frame via a promptable segmentation model, the mask is propagated across the clip, then a dedicated video matting model refines the rough mask into a per-pixel alpha matte for clean compositing onto a new background.
Source: https://arxiv.org/abs/2512.11782