arXiv AI

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.

arXiv Computer Vision
Sep 11

DirectSwap: Paired, Mask-Free Video Head Swapping with Full-Reference Evaluation

The paper introduces DirectSwap, a mask‑free video head‑swapping method that leverages a newly created cross‑identity paired dataset, HeadSwapBench. By synthesizing expression‑synchronized video pairs from real footage, the authors provide frame‑aligned ground truth for full‑reference evaluation of identity, expression, pose, reconstruction fidelity, and temporal stability. DirectSwap outperforms traditional same‑identity masked reconstruction, especially when head silhouettes change, and can restore non‑head content without external segmentation.

By Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li, Fan Yang, Guangxi Liu, Xiaodan Liang
arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
Hugging Face Trending Papers
Jul 23

GroupVideo: Multi-Identity Customized Text-to-Video Generation

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.

arXiv Computer Vision
Sep 3

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

The paper introduces a zero‑shot subject‑driven video generation framework that eliminates the need for per‑subject tuning and large subject‑video datasets. It achieves this by separating identity injection—learned from subject‑image pairs—and motion‑awareness preservation—maintained with a small set of arbitrary videos, and optimizes both with stochastic switching and dropout techniques. Using CogVideoX‑5B, the method adapts a single model with only 200K subject‑image pairs and 4,000 arbitrary videos in 288 A100 GPU hours, representing roughly 1% of the compute required by previous zero‑shot baselines while preserving subject fidelity and motion quality.

By Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo
arXiv AI
Sep 25

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

WildHSR introduces a lightweight adaptation of 3D foundation models to jointly recover metric cameras, scene geometry, and persistent person identities from monocular video. By generating pseudo‑scale labels from curated web footage and fine‑tuning a Scale Readout, the method predicts metric scale directly from foundation‑model tokens. It also exploits intermediate query‑key features to associate per‑frame bodies, enabling feed‑forward reconstruction that outperforms state‑of‑the‑art optimization‑based methods on several benchmarks while running at 10.1 fps.

By Jerrin Bright, John Zelek
arXiv Computer Vision
Aug 24

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

CogCanvas is a new benchmark for multi-subject reference-based image generation, featuring 1,952 curated reference images of 100 celebrities, 115 objects/fashion items, and 29 real-world backgrounds. It generates 1,361 compositional prompts with 2–5 people, using a pipeline that includes DINOv2 deduplication, aesthetic filtering, and automated graph derivation for interaction and positioning. The benchmark evaluates three tasks—reference-based multi-human-object generation, text-to-image compositional generation, and reference retrieval—under a six-axis protocol, and introduces BG‑Sim and Attr‑VQA metrics to assess background fidelity and attribute binding.

By Long-Bao Nguyen, Quang-Khai Le, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le