arXiv:2606. 17417v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception.
By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami
arXiv:2607. 00247v1 Announce Type: cross Abstract: Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors.
By Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang
arXiv:2608. 19211v1 Announce Type: cross Abstract: Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content.
By Linkai Peng, Baorian Nuchged
The paper introduces a method to improve audio‑visual speech recognition by applying contrastive decoding (CD) that contrasts audio‑only with audio‑visual conditioning within the same model. It addresses the issue of a fixed CD strength by scaling the influence adaptively for each token, using reliability signals from attention dynamics and predictive divergence. Experiments on the LRS3 dataset demonstrate consistent gains in both clean and low‑SNR scenarios.
By YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
The paper presents the first music‑specific, layer‑wise empirical study of hallucination in audio‑language models, framing it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. It introduces MuseDiag, a diagnostic framework that evaluates nine models and finds universal vocal misperception, significant tonal perception differences, and identifies Audio‑Flamingo‑3 as the most stable model. The study also proposes two training‑free mitigation methods, ADD‑M and TPA, which reduce hallucination in probing but show variable effectiveness in free‑form generation, highlighting the need for multi‑paradigm evaluation.
By Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu
arXiv:2608.22236v2 Announce Type: replace-cross
Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability...
By Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang