Loading...
A multimodal LLM (SEED-X), fine-tuned as a text-compatible identity adapter, reads one or more character reference images and adjusts their expression, pose, and action to match each panel's text cues; those adapted features feed an SDXL diffusion backbone through masked cross-attention alongside per-panel character and dialogue bounding boxes, producing layout-aware, identity-consistent manga pages.
Source: https://arxiv.org/abs/2412.07589