From Masking to Merging: Rethinking SpecAugment for Efficient Audio Spectrogram Transformer
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces LAST, a Looped Audio Spectrogram Transformer that processes all tokens once and then reuses the same blocks to refine only the class token over fixed audio features, making subsequent passes inexpensive. On AudioSet, a ten‑pass LAST outperforms a twelve‑layer sequential transformer by 2.1% relative mean average precision while using 49.4% fewer parameters, 42% fewer MACs, and achieving 9.8% higher throughput. Increasing the pass count from two to ten improves accuracy with only a 1.2% increase in computation, and the model shows enhanced robustness to temporal masking and other auditory augmentations across music, environmental, and event sound classification tasks.
arXiv:2511. 20973v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
arXiv:2604. 01832v1 Announce Type: cross Abstract: We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge.
arXiv:2608.30927v1 Announce Type: cross Abstract: Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language...
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.