arXiv AI

Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning

arXiv:2603. 10725v3 Announce Type: replace-cross Abstract: The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information.

arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja
arXiv AI
Sep 4

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

ToolDF is a tool‑integrated reasoning framework designed for detecting mixed‑authenticity audio deepfakes, where genuine and manipulated audio cues coexist across time or overlapping sources. It uses an audio large language model to orchestrate tasks such as source separation and routing to domain‑specific experts, aggregating their evidence into an interpretable verdict. The authors also introduce a mixed‑authenticity ADD benchmark and report that ToolDF outperforms monolithic baselines, achieving significant macro‑F1 gains while localizing evidence to specific temporal regions and acoustic sources.

By Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim
arXiv AI
Sep 12

A Survey of Threats Against Voice Authentication and Anti-Spoofing Systems

The paper reviews how voice authentication has evolved from handcrafted acoustic features to deep learning speaker embeddings, expanding its use in finance, smart devices, and law enforcement. It surveys modern threats—including data poisoning, adversarial, deepfake, and adversarial spoofing attacks—tracing their development alongside technological advances. For each attack type, the authors summarize methods, datasets, performance, and limitations, and organize the literature using accepted taxonomies to highlight emerging risks and open challenges.

By Kamel Kamel, Keshav Sood, Hridoy Sankar Dutta, Sunil Aryal
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal
arXiv AI
Sep 12

Deep-Fake CAPTCHA: Mitigating Next-Generation Social Engineering Attacks

The paper introduces DF‑CAPTCHA, an active defense that asks callers to complete simple challenge‑response tasks during voice and video calls. By evaluating responses on realism, identity consistency, task completion, and response time, the system can detect real‑time deepfake impersonations. Experiments across audio and video modalities show that DF‑CAPTCHA outperforms passive artifact‑search methods, achieving high accuracy and demonstrating the effectiveness of challenge‑based verification against next‑generation social engineering attacks.

By Guy Frankovits, Lior Yasur, Fred M. Grabovski, Yisroel Mirsky
arXiv AI
Jun 10

What Do Deepfake Speech Detectors Actually Hear?

arXiv:2606. 10912v1 Announce Type: cross Abstract: Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision.

By Vojt\v{e}ch Stan\v{e}k, Veronika Jirmusov\'a, Anton Firc, Kamil Malinka, Jakub Re\v{s}, Martin Pere\v{s}\'ini