arXiv AI

What Do Deepfake Speech Detectors Actually Hear?

arXiv:2606. 10912v1 Announce Type: cross Abstract: Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision.

arXiv Machine Learning
Sep 14

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

The paper introduces STAG, a post‑hoc framework that provides token‑level spectro‑temporal grounding for captions produced by audio‑based multimodal large language models (MLLMs). STAG estimates temporal support for each token via vocabulary projections of encoded audio, measures frequency‑band relevance through controlled spectral occlusion, and fuses these signals into a spectro‑temporal relevance map. Evaluations across ten explanation methods and four grounding benchmarks show that STAG achieves superior event‑localization performance on every dataset, and counterfactual deletion experiments confirm that removing the identified evidence selectively reduces model confidence and often eliminates the corresponding event from regenerated captions.

By Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone
arXiv AI
Sep 4

ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

ToolDF is a tool‑integrated reasoning framework designed for detecting mixed‑authenticity audio deepfakes, where genuine and manipulated audio cues coexist across time or overlapping sources. It uses an audio large language model to orchestrate tasks such as source separation and routing to domain‑specific experts, aggregating their evidence into an interpretable verdict. The authors also introduce a mixed‑authenticity ADD benchmark and report that ToolDF outperforms monolithic baselines, achieving significant macro‑F1 gains while localizing evidence to specific temporal regions and acoustic sources.

By Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim
arXiv AI
Jul 7

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.

By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
Hugging Face Trending Papers
Jun 24

From Sounds to Scenes: A Benchmark for Evaluating Context-Aware Auditory Scene Understanding in Large Audio Language Models

Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes.