arXiv Machine Learning By Haider Al-Tahan, Sean O'Brien, Anastasia Razdaibiedina, N. Apurva Ratan Murty

LAST: Looped Audio Spectrogram Transformer

Read the original on arXiv Machine Learning →

The paper introduces LAST, a Looped Audio Spectrogram Transformer that processes all tokens once and then reuses the same blocks to refine only the class token over fixed audio features, making subsequent passes inexpensive. On AudioSet, a ten‑pass LAST outperforms a twelve‑layer sequential transformer by 2.1% relative mean average precision while using 49.4% fewer parameters, 42% fewer MACs, and achieving 9.8% higher throughput. Increasing the pass count from two to ten improves accuracy with only a 1.2% increase in computation, and the model shows enhanced robustness to temporal masking and other auditory augmentations across music, environmental, and event sound classification tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

AudioFuse: Unified Spectral-Temporal Learning via a Hybrid ViT-1D CNN Architecture for Robust Phonocardiogram Classification

AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.

By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv Machine Learning
Sep 24

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

The paper introduces RALCT, a lightweight Convolutional Transformer that combines randomized audio augmentations, MFCCs, and log‑mel spectrograms to extract robust features for environmental sound recognition. With only about 310,000 parameters, RALCT achieves state‑of‑the‑art accuracy—over 93% on UrbanSound8K, peaking at 94.56%—making it suitable for deployment on mobile devices. The authors also develop a mobile app that integrates the model to provide real‑time safety alerts for hearing‑impaired users.

By Julia Huang