arXiv AI

Auditory Illusion Benchmark for Large Audio Language Models

The paper introduces AIB, the first benchmark of auditory perceptual illusion tasks designed for Large Audio Language Models (LALMs). It covers ten representative music, sound, and speech‑based illusory phenomena, each annotated for knowledge‑based priors, and evaluates models alongside controlled human listening studies. Results reveal that while LALMs are generally signal‑faithful on low‑level acoustic cues, some models show more human‑like responses when linguistic or musical priors are involved, yet none fully match human perceptual patterns.

arXiv AI
Sep 18

Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study

The paper presents the first music‑specific, layer‑wise empirical study of hallucination in audio‑language models, framing it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. It introduces MuseDiag, a diagnostic framework that evaluates nine models and finds universal vocal misperception, significant tonal perception differences, and identifies Audio‑Flamingo‑3 as the most stable model. The study also proposes two training‑free mitigation methods, ADD‑M and TPA, which reduce hallucination in probing but show variable effectiveness in free‑form generation, highlighting the need for multi‑paradigm evaluation.

By Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu
arXiv AI
Jun 11

RAIL: Rethinking Auditory Intelligence in Large Audio-Language Models with a CHC-Grounded Benchmark

arXiv:2606. 11260v1 Announce Type: cross Abstract: Humans process rich auditory environments through tightly integrated cognitive capabilities such as audio perception, audio reasoning, and memory.

By Hongyu Jin, Siyi Wang, Yang Xiao, Jiaheng Dong, Shihong Tan, Kaiyuan peng, Georgiana Juravle, Shanquan Chen, Gongping Huang, Hong Jia, Eun-Jung Holden, James Bailey, Ting Dang
arXiv Machine Learning
Jun 17

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

arXiv:2606. 17417v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami
arXiv Computation and Language
Sep 15

How Contrastive Decoding Enhances Large Audio Language Models

The paper evaluates four Contrastive Decoding (CD) strategies for Large Audio Language Models (LALMs) and finds that Audio-Aware Decoding and Audio Contrastive Decoding are the most effective. Their performance varies across models, largely depending on the baseline error profile: CD reliably fixes errors where models incorrectly claim no audio or rely on uncertainty-driven guessing, but struggles with flawed reasoning or confident misassertions. A token-level analysis shows that CD’s suppression targets hesitation markers, explaining its limited impact on confident errors.

By Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, Hung-yi Lee