arXiv Computation and Language

How Contrastive Decoding Enhances Large Audio Language Models

The paper evaluates four Contrastive Decoding (CD) strategies for Large Audio Language Models (LALMs) and finds that Audio-Aware Decoding and Audio Contrastive Decoding are the most effective. Their performance varies across models, largely depending on the baseline error profile: CD reliably fixes errors where models incorrectly claim no audio or rely on uncertainty-driven guessing, but struggles with flawed reasoning or confident misassertions. A token-level analysis shows that CD’s suppression targets hesitation markers, explaining its limited impact on confident errors.

arXiv Machine Learning
Jun 17

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

arXiv:2606. 17417v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami
arXiv Computer Vision
Aug 28

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

The paper introduces a method to improve audio‑visual speech recognition by applying contrastive decoding (CD) that contrasts audio‑only with audio‑visual conditioning within the same model. It addresses the issue of a fixed CD strength by scaling the influence adaptively for each token, using reliability signals from attention dynamics and predictive divergence. Experiments on the LRS3 dataset demonstrate consistent gains in both clean and low‑SNR scenarios.

By YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
arXiv AI
Sep 18

Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study

The paper presents the first music‑specific, layer‑wise empirical study of hallucination in audio‑language models, framing it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. It introduces MuseDiag, a diagnostic framework that evaluates nine models and finds universal vocal misperception, significant tonal perception differences, and identifies Audio‑Flamingo‑3 as the most stable model. The study also proposes two training‑free mitigation methods, ADD‑M and TPA, which reduce hallucination in probing but show variable effectiveness in free‑form generation, highlighting the need for multi‑paradigm evaluation.

By Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu
arXiv Machine Learning
Sep 2

MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models

arXiv:2608.22236v2 Announce Type: replace-cross Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability...

By Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
arXiv Computation and Language
6d ago

Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.

By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
arXiv AI
Sep 3

Auditory Illusion Benchmark for Large Audio Language Models

The paper introduces AIB, the first benchmark of auditory perceptual illusion tasks designed for Large Audio Language Models (LALMs). It covers ten representative music, sound, and speech‑based illusory phenomena, each annotated for knowledge‑based priors, and evaluates models alongside controlled human listening studies. Results reveal that while LALMs are generally signal‑faithful on low‑level acoustic cues, some models show more human‑like responses when linguistic or musical priors are involved, yet none fully match human perceptual patterns.

By Hayoon Kim, Eunice Hong, Kyogu Lee
arXiv AI
6d ago

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

The study evaluates various uncertainty estimation methods—probability-based, sampling-based, self-verification, evidential, and contrastive—across four open-weight audio-language models and five audio QA benchmarks. In multiple-choice settings, first-token probability measures outperform others, achieving a mean AUROC of .740, while open-ended evaluation shows lower accuracy but still predictive uncertainty. Ablation experiments reveal that removing audio evidence significantly degrades error-detection performance, indicating that uncertainty relies more on audio than on question text.

By Aaron Isidore Grace, Weiran Wang