WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.
The paper introduces a benchmark and evaluation system for measuring how well generative image models preserve the identity of a subject across generation, editing, restoration, and multi‑subject scenarios. It compares three paradigms—input context, trainable subject‑specific parameters, and a persistent identity layer—showing that persistent identity consistently improves fidelity while keeping image quality and instruction adherence high. The study finds that identity preservation remains a distinct limitation of current foundation models, especially under iterative edits, small scales, severe degradation, and multi‑subject composition.
arXiv:2608. 20336v1 Announce Type: new Abstract: Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people.
arXiv:2503.09130v2 Announce Type: replace-cross Abstract: This paper presents InteractEdit, a novel framework for reference-free Human-Object Interaction (HOI) editing that tackles the challenging ta...
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces.
CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.
arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.
arXiv:2609.37198v1 Announce Type: new Abstract: Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few r...
arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.
arXiv:2606. 19103v1 Announce Type: cross Abstract: Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions.
arXiv:2601.01352v2 Announce Type: replace Abstract: Human identity-preserving text-to-video generation remains challenging under large changes in viewpoint, facial expression, illumination, and motio...
arXiv:2603. 08090v3 Announce Type: replace-cross Abstract: Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting target subjects according to user instructions.
Generating and editing a person's face demands high precision, as even minor modifications can significantly alter a subject's perceived identity. Current personalization and editing methods built on general-purpose text-to-image models, however, often lack the precision required for fine-grained facial edits.
arXiv:2610.01969v1 Announce Type: new Abstract: Concept erasure aims to remove a target concept, such as a copyrighted style, a recognizable character, or unsafe content, from a pretrained text-to-im...