arXiv AI By Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu

Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study

Read the original on arXiv AI →

The paper presents the first music‑specific, layer‑wise empirical study of hallucination in audio‑language models, framing it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. It introduces MuseDiag, a diagnostic framework that evaluates nine models and finds universal vocal misperception, significant tonal perception differences, and identifies Audio‑Flamingo‑3 as the most stable model. The study also proposes two training‑free mitigation methods, ADD‑M and TPA, which reduce hallucination in probing but show variable effectiveness in free‑form generation, highlighting the need for multi‑paradigm evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 3

Auditory Illusion Benchmark for Large Audio Language Models

The paper introduces AIB, the first benchmark of auditory perceptual illusion tasks designed for Large Audio Language Models (LALMs). It covers ten representative music, sound, and speech‑based illusory phenomena, each annotated for knowledge‑based priors, and evaluates models alongside controlled human listening studies. Results reveal that while LALMs are generally signal‑faithful on low‑level acoustic cues, some models show more human‑like responses when linguistic or musical priors are involved, yet none fully match human perceptual patterns.

By Hayoon Kim, Eunice Hong, Kyogu Lee
arXiv Computation and Language
Sep 15

How Contrastive Decoding Enhances Large Audio Language Models

The paper evaluates four Contrastive Decoding (CD) strategies for Large Audio Language Models (LALMs) and finds that Audio-Aware Decoding and Audio Contrastive Decoding are the most effective. Their performance varies across models, largely depending on the baseline error profile: CD reliably fixes errors where models incorrectly claim no audio or rely on uncertainty-driven guessing, but struggles with flawed reasoning or confident misassertions. A token-level analysis shows that CD’s suppression targets hesitation markers, explaining its limited impact on confident errors.

By Tzu-Quan Lin, Wei-Ping Huang, Yi-Cheng Lin, Hung-yi Lee