Audio deepfake detectors need to transfer to languages absent from training, as multilingual speech synthesis outpaces labeled anti-spoofing resources. While detectors increasingly rely on self-superv...
arXiv:2609.13842v1 Announce Type: cross
Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing a...
By Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki
The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.
By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park
arXiv:2606. 08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning.
By Moshe Mandel, Shlomo E. Chazan
The paper investigates whether the Voxtral audio‑language model can detect speech spoofing. It shows that without task‑specific adaptation, the model’s language‑model layers prioritize semantic content, making spoof‑discriminative acoustic cues less separable. By applying lightweight weight‑decomposed low‑rank adaptation (DoRA), the authors create Spooftral, which achieves an equal error rate of 4.25% on the ASVspoof5 evaluation set.
By Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot
arXiv:2607. 11706v1 Announce Type: cross Abstract: Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks.
By Aastha Sharma, Guangjing Wang