arXiv Machine Learning

Representation Matters in Randomized Smoothing for Audio Classification

arXiv:2606. 04210v1 Announce Type: cross Abstract: Randomized smoothing (RS) certifies robustness in the vector space where Gaussian noise is added.

arXiv AI
Sep 2

MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries

MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.

By Utsab Ghosh, Roshni Chakraborty
arXiv AI
Sep 24

SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

SsgCaps is a publicly available dataset of human-engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action-based typology, enabling extensive yet plausible sampling. A comparative quantitative analysis shows only small differences between the open and private versions, supporting the recommendation of the open version for benchmarking sound scene generation algorithms.

By Modan Tailleur (LS2N), Junwon Lee (LS2N), Laurie M Heller (LS2N), Mathieu Lagrange (LS2N), Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto
arXiv Computation and Language
Sep 24

Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

Mizar is a 159.3‑million‑parameter audio‑language model designed for devices with limited memory and computation. It couples a compact CED‑Small audio encoder with SmolLM2‑135M via a frequency‑merging mapper and is trained in three stages—audio‑language alignment, audio‑dependent fine‑tuning, and post‑training—to improve performance on audio‑question tasks. Across five random seeds, Mizar outperforms all other sub‑200M‑parameter ALMs on MMAU, MMAR, and ADQA‑clean, achieving mean accuracies of 52.92%, 42.42%, and 36.02% respectively, while enabling local inference on a single CPU with an average latency of 1.09 seconds for MMAU questions.

By Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji
arXiv Machine Learning
Sep 10

Clean Accuracy Does Not Guarantee Provenance Robustness: A Prospective Codec-Stress Evaluation of Audio Attribution

The study evaluates audio provenance attribution systems, showing that high clean‑benchmark accuracy does not translate to robustness after codec compression. Using a prospectively registered protocol, the authors measured closed‑set attribution performance on two corpora after single‑stage codec transport, finding significant degradation—up to 70.3 Macro‑F1 points for WavLM‑Base+ and 61.0 for W2V2‑BERT 2.0—depending on codec settings and representation. The results demonstrate that clean accuracy alone cannot guarantee deployment robustness across different codecs and representations.

By Gang Shi (Independent Researcher)
Hugging Face Trending Papers
Aug 3

Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories.