arXiv:2606. 06357v1 Announce Type: cross Abstract: Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable.
By Dinghao Zhou, Xingchen Song, Di Wu, Pengyu Cheng, Shengfan Shen, Sixiang Lv
arXiv:2606. 11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST).
By Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
arXiv:2511. 21325v2 Announce Type: replace-cross Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs.
By Ido Nitzan Hidekel, Gal lifshitz, Khen Cohen, Dan Raviv
arXiv:2607. 00720v1 Announce Type: cross Abstract: Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time series data remains a critical yet unresolved challenge.
By Seung Hun Han, Hyeongwon Kang, Jinwoo Park, Pilsung Kang
The paper presents an autoencoder-based acoustic anomaly detection system implemented on Intel’s Loihi 2 neuromorphic processor. It achieves high detection performance—0.9959 AUC on a clean ToyADMOS ToyCar benchmark and 0.7990 source AUC on a noisy DCASE 2026 Task 2 ToyCar benchmark—while operating with only 0.0406–0.0426 mJ of dynamic energy per sample, far below CPU and GPU baselines. The system demonstrates that low‑power, on‑chip inference is feasible for persistent machine monitoring.
By Steven C. Nesbit (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA), Victor M. Vergara (AeroVironment Inc., Albuquerque, USA), Michael A. Felix (University of New Mexico COSMIAC Research Center, Albuquerque, USA), Evan T. Kain (Air Force Research Laboratory, Kirtland AFB, USA), Luis R. Garc\'ia Carrillo (Air Force Research Laboratory, Kirtland AFB, USA), Gerd J. Kunde (Nuclear and Particle Physics and Applications, P-3, Los Alamos National Laboratory, Los Alamos, USA), Andrew T. Sornborger (Information Sciences, CAI-3, Los Alamos National Laboratory, Los Alamos, USA)
arXiv:2607. 17761v1 Announce Type: cross Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress.
By Jun Xue, Zhuolin Yi, Yanzhen Ren, Yihuan Huang, Jiayu Xiong, Yi Chai, Guanxiang Feng, Jiajun Liu, Tong Zhang
arXiv:2608. 05705v1 Announce Type: cross Abstract: Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings.
By Victor Gialis, Maxime Metz, David Esteve, Abdenour Soualhi
arXiv:2509. 18751v4 Announce Type: replace Abstract: Recently reconstruction-based deep models have been widely used for time series anomaly detection, but as their capacity and generalization capability increase, these models tend to over-generalize, often reconstructing unseen anomalies accurately.
By Samuel Yoon, Jongwon Kim, Juyoung Ha, Young Myoung Ko
arXiv:2606. 06907v1 Announce Type: cross Abstract: Large audio language models (LALMs) extend large language models with an audio encoder and large-scale audio data.
By Seonuk Kim, Yonghyeon Jun, Ju Yeon Kang, Jimin Hong, Yoonhyeong Lee, Nam Soo Kim
arXiv:2608. 15037v1 Announce Type: cross Abstract: Audio-Text Foundation Models (ATMs) fail catastrophically under severe acoustic noise, yet existing adaptation strategies either rely on gradient-based Test-Time Adaptation (TTA), which reinforces noise rather than signal, or on prompt tuning that requires privileged noise annotations unavailable at inference.
By Ashish Anand Shukla, Rini Smita Thakur, Aryan Das, Vinod K. Kurmi
arXiv:2608. 09158v1 Announce Type: cross Abstract: Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs.
By Yuanhe Zhang, Weiliu Wang, Jie Ren, Liang Lin, Zhenhong Zhou, Haoran Gao, Kun Wang, Chen Li, Li Sun, Sen Su
arXiv:2606. 30646v1 Announce Type: cross Abstract: Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs, providing a non-invasive proxy for cognitive assessment.
By Chukwuemeka Ugwu, Oluwafemi Richard Oyeleke