arXiv AI By Aaron Isidore Grace, Weiran Wang

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

Read the original on arXiv AI →

The study evaluates various uncertainty estimation methods—probability-based, sampling-based, self-verification, evidential, and contrastive—across four open-weight audio-language models and five audio QA benchmarks. In multiple-choice settings, first-token probability measures outperform others, achieving a mean AUROC of .740, while open-ended evaluation shows lower accuracy but still predictive uncertainty. Ablation experiments reveal that removing audio evidence significantly degrades error-detection performance, indicating that uncertainty relies more on audio than on question text.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 30

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

arXiv:2512. 09066v2 Announce Type: replace-cross Abstract: Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art.

By \v{S}imon Sedl\'a\v{c}ek, Sara Barahona, Bolaji Yusuf, Laura Herrera-Alarc\'on, Santosh Kesiraju, Cecilia Bola\~nos, Alicia Lozano-Diez, Sathvik Udupa, Fernando L\'opez, Allison Ferner, Ramani Duraiswami, Jan \v{C}ernock\'y
arXiv Computation and Language
Sep 7

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

The paper introduces pseudo‑ensembles for music audio‑language models to enable abstention when the model is uncertain. By perturbing inputs—such as shuffling answer order, corrupting audio, or swapping option labels—multiple predictive distributions are generated from a single pretrained model, allowing the use of ensemble‑based uncertainty metrics like entropy, expected entropy, and mutual information. Experiments on TinyMU with MuChoMusic show that averaging over four answer orderings improves accuracy from 55.7% to 59.2% and reduces the error‑retention curve area from 0.293 to 0.261, all with only a few extra forward passes and no retraining.

By Aanya Maheshwari, Vatsal Raina
arXiv AI
3d ago

Audio LLMs Know When They Can't Hear You

The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.

By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh