arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.
By Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.
arXiv:2610.11023v1 Announce Type: new
Abstract: Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait capture...
By Tianwen Fu, Wenbin Teng, Gonglin Chen, Junyi Ouyang, Haolin Xiong, Yajie Zhao
arXiv:2601.01352v2 Announce Type: replace
Abstract: Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motio...
By Yixuan Lai, He Wang, Kun Zhou, Tianjia Shao
arXiv:2605. 02814v2 Announce Type: replace-cross Abstract: Severe face degradation can remove person-specific evidence, making restoration underdetermined.
By Axi Niu, Jinyang Zhang, Senyan Qing
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.
By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le