arXiv:2607. 04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set.
By Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
arXiv:2606. 16137v1 Announce Type: cross Abstract: Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making.
By Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang, Berrak Sisman, Bj\"orn W. Schuller
arXiv:2607. 12584v1 Announce Type: cross Abstract: The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics.
By Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
Recent Large Audio Language Models (LALMs) have achieved remarkable progress in audio perceptual tasks across individual acoustic layers, including speech, sound, and music. However, existing benchmarks predominantly evaluate these layers in isolation, overlooking the complex contextual relationships that arise when multiple acoustic sources co-occur in real-world auditory scenes.
arXiv:2606. 14466v1 Announce Type: cross Abstract: This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection.
By Piotr Kit{\l}owski, Dominik Wi\k{a}cek, Mateusz Modrzejewski
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity.