arXiv:2609.13260v1 Announce Type: cross
Abstract: This paper proposes SpecAugment-Patch Merging, a simple yet effective method to accelerate Audio Spectrogram Transformer (AST) training. We first app...
By Minhee Park, Hyowon Ahn, Chanwoo Kim
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.
By Julia Huang
arXiv:2606. 11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST).
By Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
arXiv:2606. 06907v1 Announce Type: cross Abstract: Large audio language models (LALMs) extend large language models with an audio encoder and large-scale audio data.
By Seonuk Kim, Yonghyeon Jun, Ju Yeon Kang, Jimin Hong, Yoonhyeong Lee, Nam Soo Kim
arXiv:2511. 13487v3 Announce Type: replace-cross Abstract: This study presents a systematic evaluation of time-frequency feature design for binaural sound source localization (SSL), focusing on how feature selection influences model performance across diverse conditions.
By Davoud Shariat Panah, Alessandro Ragano, Dan Barry, Jan Skoglund, Andrew Hines
UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.
By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
arXiv:2606. 00081v1 Announce Type: cross Abstract: Distributed Acoustic Sensing (DAS) enables large-scale monitoring through optical fibers, but its high dimensionality and complex spatio-temporal patterns make event classification demanding.
By Michel Dione (CERI SN - IMT Nord Europe), Jerry Lonlac (CERI SN - IMT Nord Europe), H\'el\`ene Louis (CERI SN - IMT Nord Europe), Anthony Fleury (CERI SN - IMT Nord Europe), Stephane Lecoeuche
arXiv:2606. 17301v1 Announce Type: cross Abstract: Search, a foundational operation in computer science, maps a query to a matching item in a collection.
By Muhammad Taimoor Haseeb, Ahmad Hammoudeh, Gus Xia
arXiv:2607. 06179v1 Announce Type: cross Abstract: There are some datasets of varying scales for audio classification (AC) applied to different tasks.
By Hong Lyu, Mingru Yang, Qianhua He, Yanxiong Li, Jinxin Huang, Zhengyu Pei
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
By Fengrui Liu, Ruiyang Huang, Qijian Zheng, Yuanfang Wang, Feng Liu
arXiv:2511. 20973v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.
By Saurabhchand Bhati, Samuel Thomas, Hilde Kuehne, Rogerio Feris, James Glass