arXiv AI

Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors

arXiv:2606. 29544v1 Announce Type: cross Abstract: We present Proteus, a framework developed at Resemble AI for automated robustness testing of our audio deepfake detection system.

arXiv AI
Sep 18

Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection

The paper introduces ROGUE, a framework that builds robust audio deepfake detection workflows by combining multiple detection tools. ROGUE treats workflow creation as a sequential decision problem and uses a dual-agent system: a perturbation agent generates audio distortions while a policy agent selects and executes detection tools that are resilient to those perturbations. Experiments on several datasets and real-world corruptions show that ROGUE consistently outperforms strong baselines in robustness and generalization, demonstrating the value of adversarially optimized workflow generation for reliable deployment.

By Xiang Li, Pin-Yu Chen, Wenqi Wei
arXiv AI
Aug 17

Teffic-Audio: Tell Fact from Fiction

arXiv:2607. 28351v2 Announce Type: replace-cross Abstract: Speech deepfake detection has expanded in scope with increasingly heterogeneous spoofing mechanisms, including speech synthesis, voice conversion, vocoder reconstruction, and neural-codec resynthesis.

By Wan Lin, Li Wang, Jindong Wang, Kunyu Feng, Zhizheng Wu
arXiv AI
Aug 11

MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection

arXiv:2608. 09593v1 Announce Type: cross Abstract: Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enabling a realistic attack scenario in which speech and background audio are independently manipulated over otherwise authentic video.

By Yanqiu Li, Yang Xiao, Jisheng Bai, Bin Chen, Hong Jia, Ting Dang
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
arXiv AI
Sep 12

A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems

The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.

By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv AI
Sep 21

A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

The paper introduces a training‑free proactive defense for detecting partial deepfake speech by using self‑embedding steganography. It embeds a compressed version of the clean audio within itself, allowing post‑hoc extraction of reference content and enabling detection of spoofed segments via codec‑based restoration. Experiments on a benchmark dataset show that this method complements passive detectors and operates without any training, offering a robust, data‑efficient alternative for partial deepfake detection.

By Yigitcan \"Ozer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi
arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja