arXiv:2606. 16532v1 Announce Type: cross Abstract: Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage.
By Zhuodong Liu, Hugen Lv, Xiangyu Li, Chunhong Yuan
The paper proposes a method called language orthogonalization to improve zero‑shot cross‑lingual audio deepfake detection. By removing language‑dependent variation from self‑supervised speech models using a target‑free ridge map on language‑identification embeddings, the approach consistently lowers equal error rates across six languages and six model backbones. The gains are larger when the target language is more distant in the language‑identification space.
By Minu Kim, Ji Sub Um, Hoirin Kim
arXiv:2609.38887v1 Announce Type: cross
Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
By Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna
arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.
By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-superv...
arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.
By Karl El Hajal, Mathew Magimai. -Doss
arXiv:2607. 03134v1 Announce Type: cross Abstract: Recent research expands beyond binary anti-spoofing with the emergence of Source Tracing, the task of identifying the specific generative origins of synthetic speech.
By Santiago Rubio, Antonio Almud\'evar, Antonio Miguel, Eduardo Lleida, Alfonso Ortega
arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.
By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
arXiv:2603. 21875v2 Announce Type: replace-cross Abstract: Speech deepfake source verification systems aims to determine whether two synthetic speech utterances originate from the same source generator, often assuming that the resulting source embeddings are independent of speaker traits.
By Xi Xuan, Wenxin Zhang, Zhiyu Li, Jennifer Williams, Ville Hautam\"aki, Tomi H. Kinnunen
arXiv:2609.26648v1 Announce Type: cross
Abstract: Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under chal...
By Pu Wang, Yujun Wang, Hugo Van hamme
arXiv:2407.04291v4 Announce Type: replace-cross
Abstract: Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech sys...
By Ismail Rasim Ulgen, John H. L. Hansen, Carlos Busso, Berrak Sisman
arXiv:2603.07551v3 Announce Type: replace-cross
Abstract: Recent zero-shot Text-to-Speech (TTS) systems can clone previously unseen voices from only a few seconds of audio. We formulate Speech Genera...
By Thanathai Lertpetchpun, Thanapat Trachu, Sai Praneeth Karimireddy, Shrikanth Narayanan