arXiv AI

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

The study evaluates various uncertainty estimation methods—probability-based, sampling-based, self-verification, evidential, and contrastive—across four open-weight audio-language models and five audio QA benchmarks. In multiple-choice settings, first-token probability measures outperform others, achieving a mean AUROC of .740, while open-ended evaluation shows lower accuracy but still predictive uncertainty. Ablation experiments reveal that removing audio evidence significantly degrades error-detection performance, indicating that uncertainty relies more on audio than on question text.

arXiv AI
Jun 30

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering

arXiv:2512. 09066v2 Announce Type: replace-cross Abstract: Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art.

By \v{S}imon Sedl\'a\v{c}ek, Sara Barahona, Bolaji Yusuf, Laura Herrera-Alarc\'on, Santosh Kesiraju, Cecilia Bola\~nos, Alicia Lozano-Diez, Sathvik Udupa, Fernando L\'opez, Allison Ferner, Ramani Duraiswami, Jan \v{C}ernock\'y
arXiv Computation and Language
Sep 7

Knowing When Not to Answer: Pseudo-Ensembles for Abstention in Music Audio-Language Models

The paper introduces pseudo‑ensembles for music audio‑language models to enable abstention when the model is uncertain. By perturbing inputs—such as shuffling answer order, corrupting audio, or swapping option labels—multiple predictive distributions are generated from a single pretrained model, allowing the use of ensemble‑based uncertainty metrics like entropy, expected entropy, and mutual information. Experiments on TinyMU with MuChoMusic show that averaging over four answer orderings improves accuracy from 55.7% to 59.2% and reduces the error‑retention curve area from 0.293 to 0.261, all with only a few extra forward passes and no retraining.

By Aanya Maheshwari, Vatsal Raina
arXiv AI
3d ago

Audio LLMs Know When They Can't Hear You

The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.

By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
arXiv Machine Learning
Sep 23

Unified Multimodal Uncertain Inference

Unified Multimodal Uncertain Inference (UMUI) is a new task that requires models to generate calibrated probability estimates for hypotheses conditioned on premises across text, audio, and video modalities. The authors create a human‑annotated evaluation set with scalar probability judgments for audio, visual, and audiovisual settings, and benchmark their approach on existing text and audio datasets. Their CLUE framework, which blends self‑consistent teacher calibration with distribution‑based confidence probing, enables a 3B‑parameter model to match or surpass zero‑shot baselines up to 32B parameters across all modalities.

By Dengjia Zhang, Alexander Martin, William Jurayj, Kenton Murray, Benjamin Van Durme, Reno Kriz
arXiv Computation and Language
Sep 15

How Contrastive Decoding Enhances Large Audio Language Models

The paper evaluates four Contrastive Decoding (CD) strategies for Large Audio Language Models (LALMs) and finds that Audio-Aware Decoding and Audio Contrastive Decoding are the most effective. Their performance varies across models, largely depending on the baseline error profile: CD reliably fixes errors where models incorrectly claim no audio or rely on uncertainty-driven guessing, but struggles with flawed reasoning or confident misassertions. A token-level analysis shows that CD’s suppression targets hesitation markers, explaining its limited impact on confident errors.

By Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, Hung-yi Lee
arXiv Machine Learning
Jun 30

RA-QA: A Benchmarking System for Respiratory Audio Question Answering Under Real-World Heterogeneity

arXiv:2602. 18452v3 Announce Type: replace-cross Abstract: As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions.

By Gaia A. Bertolino, Yuwei Zhang, Tong Xia, Domenico Talia, Cecilia Mascolo
arXiv AI
Sep 17

Label-Confidence-Aware Uncertainty Estimation in Natural Language Generation

The paper introduces Label-Confidence-Aware Uncertainty Quantification (LCA-UQ), a method that uses Pointwise Kullback-Leibler divergence to align global entropy from multiple stochastic samples with the local confidence of a candidate answer. By bridging this gap, LCA-UQ improves the reliability and stability of uncertainty assessments in natural language generation. Experiments on popular LLMs and NLP datasets show that label sources significantly influence classification and that LCA-UQ outperforms existing uncertainty estimation approaches.

By Qinhong Lin, Yinglun Feng, Yuhao Zhang, Zhongliang Yang, Linna Zhou