arXiv AI By Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

Read the original on arXiv AI →

arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 11

DirectSwap: Paired, Mask-Free Video Head Swapping with Full-Reference Evaluation

The paper introduces DirectSwap, a mask‑free video head‑swapping method that leverages a newly created cross‑identity paired dataset, HeadSwapBench. By synthesizing expression‑synchronized video pairs from real footage, the authors provide frame‑aligned ground truth for full‑reference evaluation of identity, expression, pose, reconstruction fidelity, and temporal stability. DirectSwap outperforms traditional same‑identity masked reconstruction, especially when head silhouettes change, and can restore non‑head content without external segmentation.

By Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li, Fan Yang, Guangxi Liu, Xiaodan Liang
arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
Hugging Face Trending Papers
Jul 23

GroupVideo: Multi-Identity Customized Text-to-Video Generation

Current identity customized video generation methodologies are predominantly limited to single-identity scenarios, as the lack of explicit identity separation mechanisms often leads to identity confusion in multi-identity settings. Existing multi-identity approaches, which directly extend single-identity frameworks by concatenating face images as input conditions, frequently result in unnatural facial expressions and motions, manifesting as the "copy-paste" phenomenon.

arXiv Computer Vision
Sep 3

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute

The paper introduces a zero‑shot subject‑driven video generation framework that eliminates the need for per‑subject tuning and large subject‑video datasets. It achieves this by separating identity injection—learned from subject‑image pairs—and motion‑awareness preservation—maintained with a small set of arbitrary videos, and optimizes both with stochastic switching and dropout techniques. Using CogVideoX‑5B, the method adapts a single model with only 200K subject‑image pairs and 4,000 arbitrary videos in 288 A100 GPU hours, representing roughly 1% of the compute required by previous zero‑shot baselines while preserving subject fidelity and motion quality.

By Daneul Kim, Jingxu Zhang, Wonjoon Jin, Sunghyun Cho, Qi Dai, Jaesik Park, Chong Luo