arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
arXiv:2608. 20326v1 Announce Type: cross Abstract: Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions.
By Parampreet Singh, Anushka Singh, Sumit Kumar, Vipul Arora
The study evaluates various uncertainty estimation methods—probability-based, sampling-based, self-verification, evidential, and contrastive—across four open-weight audio-language models and five audio QA benchmarks. In multiple-choice settings, first-token probability measures outperform others, achieving a mean AUROC of .740, while open-ended evaluation shows lower accuracy but still predictive uncertainty. Ablation experiments reveal that removing audio evidence significantly degrades error-detection performance, indicating that uncertainty relies more on audio than on question text.
By Aaron Isidore Grace, Weiran Wang
Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. Consequently, users lack a reliable signal for deciding when a prediction can be trusted.
The paper investigates whether audio large language models (Audio LLMs) can detect when their own transcriptions are unreliable. It finds that the models are poor at self-assessment and that existing methods offer limited detection. By leveraging audio-encoder representations, the authors develop a lightweight predictor that accurately flags unreliable transcriptions and can prompt user clarification without altering the underlying model.
By Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh
We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories.