arXiv Computer Vision By Yanan Wang, Shengcai Liao, Panwen Hu, Xin Li, Fan Yang, Guangxi Liu, Xiaodan Liang

DirectSwap: Paired, Mask-Free Video Head Swapping with Full-Reference Evaluation

Read the original on arXiv Computer Vision →

The paper introduces DirectSwap, a mask‑free video head‑swapping method that leverages a newly created cross‑identity paired dataset, HeadSwapBench. By synthesizing expression‑synchronized video pairs from real footage, the authors provide frame‑aligned ground truth for full‑reference evaluation of identity, expression, pose, reconstruction fidelity, and temporal stability. DirectSwap outperforms traditional same‑identity masked reconstruction, especially when head silhouettes change, and can restore non‑head content without external segmentation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Jun 11

ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

arXiv:2606. 11670v1 Announce Type: cross Abstract: Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references.

By Zijie Meng, Jiwen Liu, Yufei Liu, Chengzhuo Tong, Xiaoqiang Liu, Yuanxing Zhang, Yulong Xu, Pengfei Wan
arXiv Computer Vision
Aug 21

ID-V2V: Identity-Preserving Video Restylization

arXiv:2607. 22830v2 Announce Type: replace Abstract: In visual storytelling, human performances are central to creative intent and narrative meaning.

By Yuancheng Xu, Mingming He, Pablo Salamanca, Li Ma, Yash Kant, Emmett Steven, Paul Debevec, Ning Yu
arXiv Computer Vision
Aug 28

EditaLive! Unified Character Video Editing for Live Streaming

EditaLive! is a new real‑time framework for character video editing in live streaming, built on a pretrained image animation model (Wan‑Animate) that separates appearance from motion. It uses the CharEdit‑50K dataset for reference‑frame editing and video reconstruction, and adapts the model from offline bidirectional to causal streaming generation. A self‑rollout distillation strategy compresses the model into a two‑step sampler, employing fixed RoPE, alignment forcing, and first‑frame preserved sparse attention to reduce appearance drift and achieve low‑latency inference while preserving facial expressions.

By Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang, Bo Li, Xiaodong Cun