WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces.
arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.
Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.
arXiv:2610.11023v1 Announce Type: new Abstract: Identity-preserving video generation aims to maintain a subject's identity while synthesizing realistic videos. Yet a single reference portrait capture...
arXiv:2601.01352v2 Announce Type: replace Abstract: Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motio...
arXiv:2605. 02814v2 Announce Type: replace-cross Abstract: Severe face degradation can remove person-specific evidence, making restoration underdetermined.
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.
arXiv:2511.05575v2 Announce Type: replace Abstract: Diffusion-based approaches have recently achieved strong results in face swapping, offering improved visual quality over traditional GAN-based meth...
The paper introduces a benchmark and evaluation system for measuring how well generative image models preserve the identity of a subject across generation, editing, restoration, and multi‑subject scenarios. It compares three paradigms—input context, trainable subject‑specific parameters, and a persistent identity layer—showing that persistent identity consistently improves fidelity while keeping image quality and instruction adherence high. The study finds that identity preservation remains a distinct limitation of current foundation models, especially under iterative edits, small scales, severe degradation, and multi‑subject composition.
arXiv:2607. 16287v1 Announce Type: cross Abstract: Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended to reconstruct and animate 3D face models from a small number of images.
arXiv:2608.23410v1 Announce Type: new Abstract: Photorealistic novel view synthesis of people remains challenging at high spatial resolutions and across multiple target cameras, where preserving iden...