Loading...
Tencent's HunyuanCustom fuses a LLaVA-based text-image understanding module with the HunyuanVideo DiT and an image-ID enhancement module (temporal concatenation) to keep one or more reference subjects visually consistent throughout a generated or subject-replaced video.
Source: https://arxiv.org/abs/2505.04512