The paper proposes a method called language orthogonalization to improve zero‑shot cross‑lingual audio deepfake detection. By removing language‑dependent variation from self‑supervised speech models using a target‑free ridge map on language‑identification embeddings, the approach consistently lowers equal error rates across six languages and six model backbones. The gains are larger when the target language is more distant in the language‑identification space.
By Minu Kim, Ji Sub Um, Hoirin Kim
arXiv:2609.13842v1 Announce Type: cross
Abstract: Recent advances in speech synthesis and voice conversion have made deepfake speech increasingly realistic, making generalization to unseen spoofing a...
By Minh-Xuan Phan, Khalid Zaman, Candy Olivia Mawalim, Masashi Unoki
The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.
By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park
arXiv:2606. 08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning.
By Moshe Mandel, Shlomo E. Chazan
The paper investigates whether the Voxtral audio‑language model can detect speech spoofing. It shows that without task‑specific adaptation, the model’s language‑model layers prioritize semantic content, making spoof‑discriminative acoustic cues less separable. By applying lightweight weight‑decomposed low‑rank adaptation (DoRA), the authors create Spooftral, which achieves an equal error rate of 4.25% on the ASVspoof5 evaluation set.
By Avishai Weizman, Yehuda Ben-Shimol, Itshak Lapidot
arXiv:2608. 15037v1 Announce Type: cross Abstract: Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference.
By Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2606. 16837v1 Announce Type: cross Abstract: Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation.
By Mahtab Masoudi Nezhad, Nima Karimian
arXiv:2608. 13817v1 Announce Type: cross Abstract: Human speech production is constrained by physiology, giving rise to characteristic temporal structure on acoustic signals.
By Tom\'as Andrade Weber
arXiv:2606. 16532v1 Announce Type: cross Abstract: Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts, known as implicit identity leakage.
By Zhuodong Liu, Hugen Lv, Xiangyu Li, Chunhong Yuan
Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignment exists between the encoder and the LLM.
arXiv:2607. 11706v1 Announce Type: cross Abstract: Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represented in many legacy spoofing benchmarks.
By Aastha Sharma, Guangjing Wang