arXiv:2610.08417v1 Announce Type: new
Abstract: Recent advances in LipSync generation technology have led to the creation of highly realistic videos, posing severe societal risks. However, existing d...
By Tianyi She, Jiawei Liu, Weifeng Liu, Hanqing Zhao, Weiming Zhang, Kejiang Chen
arXiv:2605. 27944v2 Announce Type: replace Abstract: With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical.
By Ke Liu, Jiwei Wei, Wenyu Zhang, Shuchang Zhou, Ruikun Chai, Yutao Dai, Chaoning Zhang, Yang Yang
arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.
By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv:2609.38019v1 Announce Type: new
Abstract: We present RGOR (Reference-Grounded Oral Refinement), an audio-driven lip-sync framework that renders the mouth of the specific person being dubbed rat...
By Bangxun Tang
The paper introduces an end‑to‑end framework for personalized audio‑driven facial motion that operates in real time without look‑ahead. It combines a causal multi‑resolution motion tokenizer, which captures both global temporal context and fine articulatory details, with a multi‑modal style retriever that pulls stylistic priors from arbitrary reference footage using ongoing audio and motion queries. This approach allows high‑fidelity, identity‑consistent animation from just a few casually recorded clips, outperforming existing methods in lip‑sync accuracy, identity consistency, and perceived realism while maintaining real‑time streaming constraints.
By Xuangeng Chu, Yu Han, Wei Mao, Shih-En Wei
arXiv:2609.12668v1 Announce Type: new
Abstract: Recent deepfake detection studies increasingly suggest remote photoplethysmography (rPPG) signals as an authenticity cue. However, existing benchmarks...
By Chenxi Yang, Yassine Ouzar, Larbi Boubchir
arXiv:2606. 01031v1 Announce Type: cross Abstract: Audio-driven talking-head generation has advanced rapidly, yet existing evaluation protocols mainly rely on frame-wise metrics that assume strict temporal correspondence between generated and reference videos.
By Zhicheng Zhang, Lei Wang, Yu Zhang, Yongsheng Gao
ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.
By Jiaran Cai, Xingpei Ma, Shenneng Huang
Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.
By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv:2607. 25543v1 Announce Type: cross Abstract: Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes.
By Jielun Peng, Yabin Wang, Yaqi Li, Jincheng Liu, Xiaopeng Hong, Athanasios V. Vasilakos
The paper introduces Traceable TTS, a framework that enables Text‑to‑Speech systems to attribute synthesized speech to their source models without embedding explicit watermarks. By jointly training the TTS model and a discriminator, the method improves traceability generalization while maintaining or slightly enhancing audio quality. This represents the first attempt at watermark‑free TTS with strong traceability, and the authors plan to release the code to support further research.
By Yuxiang Zhao, Yunchong Xiao, Yushen Chen, Zhikang Niu, Shuai Wang, Kai Yu, Xie Chen
This paper presents an interpretable deepfake detection framework that explicitly encodes physically grounded forensic cues to analyze spatially and temporally coherent facial features in video sequences. The method transforms videos into identity-consistent facial trajectories, segments them into fixed-length temporal windows, and represents each frame with 68 structured descriptors across photometric, textural, geometric, and compression domains. These descriptors are processed by an LSTM to capture temporal dependencies, achieving strong F1-scores on four benchmark datasets and demonstrating robust cross-dataset generalization.
By Chahira Benhama, Mohand Sa\"id Allili, Assia Hamadene