AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.
By Pooneh Mousavi, Amir Ivry, Mirco Ravanelli, Cem Subakan
The paper evaluates four Contrastive Decoding (CD) strategies for Large Audio Language Models (LALMs) and finds that Audio-Aware Decoding and Audio Contrastive Decoding are the most effective. Their performance varies across models, largely depending on the baseline error profile: CD reliably fixes errors where models incorrectly claim no audio or rely on uncertainty-driven guessing, but struggles with flawed reasoning or confident misassertions. A token-level analysis shows that CD’s suppression targets hesitation markers, explaining its limited impact on confident errors.
By Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, Hung-yi Lee
arXiv:2608. 19936v1 Announce Type: cross Abstract: Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities.
By Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr C{\l}apa, Jens Madsen, Panagiotis Tzirakis
arXiv:2609.36577v1 Announce Type: cross
Abstract: Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually co...
By Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma, Lianyu Hu, Yang Liu
arXiv:2606. 17417v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception.
By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami
arXiv:2606. 10223v1 Announce Type: cross Abstract: Attributing a synthetic utterance to its originating system remains an open challenge: closed-set models fail to reject unseen synthesizers and produce overconfident predictions.
By Awais Khan, Kutub Uddin, Khalid Malik