arXiv Computer Vision By Hengyuan Xu, Qixun Wang, Yiji Cheng, Miles Yang, Zhao Zhong, Wei Cheng, Xingjun Ma, Yu-gang Jiang

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation

Read the original on arXiv Computer Vision →

arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Jul 23

GroupVideo: Multi-Identity Customized Text-to-Video Generation

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.

arXiv Computer Vision
Sep 4

Persistent Identity Preservation in Generative Image Models: A Benchmark and Evaluation System

The paper introduces a benchmark and evaluation system for measuring how well generative image models preserve the identity of a subject across generation, editing, restoration, and multi‑subject scenarios. It compares three paradigms—input context, trainable subject‑specific parameters, and a persistent identity layer—showing that persistent identity consistently improves fidelity while keeping image quality and instruction adherence high. The study finds that identity preservation remains a distinct limitation of current foundation models, especially under iterative edits, small scales, severe degradation, and multi‑subject composition.

By Mengwei Ren, Xuaner Zhang, Zhihao Xia
arXiv AI
Jun 11

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.

By Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan