arXiv Computation and Language By Aanya Maheshwari, Vatsal Raina

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

Read the original on arXiv Computation and Language →

The paper introduces pseudo‑ensembles for music audio‑language models to enable abstention when the model is uncertain. By perturbing inputs—such as shuffling answer order, corrupting audio, or swapping option labels—multiple predictive distributions are generated from a single pretrained model, allowing the use of ensemble‑based uncertainty metrics like entropy, expected entropy, and mutual information. Experiments on TinyMU with MuChoMusic show that averaging over four answer orderings improves accuracy from 55.7% to 59.2% and reduces the error‑retention curve area from 0.293 to 0.261, all with only a few extra forward passes and no retraining.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
6d ago

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

The study evaluates various uncertainty estimation methods—probability-based, sampling-based, self-verification, evidential, and contrastive—across four open-weight audio-language models and five audio QA benchmarks. In multiple-choice settings, first-token probability measures outperform others, achieving a mean AUROC of .740, while open-ended evaluation shows lower accuracy but still predictive uncertainty. Ablation experiments reveal that removing audio evidence significantly degrades error-detection performance, indicating that uncertainty relies more on audio than on question text.

By Aaron Isidore Grace, Weiran Wang
arXiv AI
3d ago

Audio LLMs Know When They Can't Hear You

The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.

By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
Hugging Face Trending Papers
Aug 3

Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories.