arXiv Computer Vision

Ariadne's Thread of LipSync: Unraveling Forgeries via Inconsistency between Lip Motions and Head Poses

arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv Computer Vision
Sep 21

Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal Retrieval

The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.

By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei
arXiv Computer Vision
Sep 25

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.

By Jiaran Cai, Xingpei Ma, Shenneng Huang
arXiv Computer Vision
Sep 23

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.

By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv Computer Vision
3d ago

Interpretable Deepfake Detection in Videos via Explicit Forensic Features and Temporal Modeling

This paper presents an interpretable deepfake detection framework that explicitly encodes physically grounded forensic cues to analyze spatially and temporally coherent facial features in video sequences. The method transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame with 68 structured descriptors across photometric, textural, geometric, and compression domains. These descriptors are processed by an LSTM to capture temporal dependencies, achieving strong F1-scores on four benchmark datasets and demonstrating robust cross-dataset generalization.

By Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene