arXiv Machine Learning

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.

arXiv AI
Jun 4

The Differentiable Auditory Loop (DAL): An ML Framework for Hyper-Personalized Hearing Aids

arXiv:2606. 04103v1 Announce Type: cross Abstract: Conventional hearing aids rely on fixed, frequency-dependent amplification and compression to manage reduced sensitivity, which often fails to provide sufficient listening support in complex environments, such as situations with multiple speakers (the ``cocktail party'' problem).

By Alejandro Ballesta Rosen, Jason Mikiel-Hunter, Julian Maclaren, Jack Collins, Richard F. Lyon, Simon Carlile
arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv AI
Aug 25

Robust Lightweight Deep Learning Models for Oral Cancer Screening

arXiv:2608.21583v1 Announce Type: new Abstract: Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While point-of-care sc...

By Siddhant Bharadwaj, Aakash Shedsale, Tejashree Subramanya, Mohd. Azfar, Praveen Birur, Debnath Pal, Shankararama Sharma, Anupama Shetty, Rajesh Sundaresan
arXiv Machine Learning
Sep 11

Single Microphone Own Voice Detection based on Simulated Transfer Functions for Hearing Aids

The paper introduces a simulation-based method for detecting a user's own voice in hearing aids using only a single microphone. It employs a data augmentation strategy with simulated acoustic transfer functions to train a transformer classifier, achieving over 90% accuracy on both simulated and real-world recordings. The approach reduces hardware complexity and power consumption while maintaining robust performance across varied spatial conditions.

By Mathuranathan Mayuravaani, W. Bastiaan Kleijn, Andrew Lensen, Charlotte S{\o}rensen
arXiv AI
Sep 2

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.

By Utsab Ghosh, Roshni Chakraborty
arXiv AI
Sep 12

Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems

The paper introduces the Spectral Masking and Interpolation Attack (SMIA), a black‑box adversarial technique that subtly alters inaudible frequency regions of AI‑generated audio to fool voice authentication systems and their countermeasures. Experiments show SMIA achieves at least 82% success against combined verification and countermeasure systems, 97.5% against standalone speaker verification, and 100% against countermeasures, revealing a critical security gap. The authors argue that current static defenses are inadequate and call for dynamic, context‑aware defenses that can adapt to evolving threats.

By Kamel Kamel, Hridoy Sankar Dutta, Keshav Sood, Sunil Aryal