arXiv Machine Learning

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

arXiv AI
Sep 1

Uncertainty Makes It Stable: Curiosity-Driven Quantized Mixture-of-Experts

The paper introduces a curiosity‑driven quantized Mixture‑of‑Experts framework that routes inputs based on Bayesian epistemic uncertainty across heterogeneous experts (BitNet ternary, 1‑16 bit BitLinear, post‑training quantization). On audio classification benchmarks, 4‑bit quantization preserves 99.9 % of full‑precision F1 while achieving 4× compression and 31 % energy savings, and curiosity‑driven routing further improves accuracy and reduces cross‑fold variance by up to 85 %. The routing is self‑organizing, allocating the most uncertain samples to the high‑precision expert, and the method demonstrates statistical parity with full precision across datasets.

By Sebasti\'an Andr\'es Cajas Ord\'o\~nez, Luis Fernando Torres Torres, Mackenzie J. Meni, Carlos Andr\'es Duran Paredes, Eric Arazo, Cristian Bosch, Ricardo Simon Carbajo, Yuan Lai, Leo Anthony Celi
arXiv Machine Learning
Sep 24

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.

By Julia Huang
arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv AI
1d ago

UniAE-MoE: A Unified Audio Encoder via Mixture of Experts

UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.

By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu