arXiv Computation and Language By Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee

Causal Tracing of Audio-Text Fusion in Large Audio Language Models

Read the original on arXiv Computation and Language →

The study applies causal tracing to large audio language models (LALMs) to uncover how they fuse acoustic and textual information. Layer‑wise analysis reveals distinct fusion strategies—progressive integration in DeSTA versus abrupt late‑stage fusion in Qwen—while token‑wise analysis identifies the final sequence token as an informational bottleneck that decisively retrieves audio content. Additionally, an attention‑like query mechanism at intermediate tokens is observed, prompting the model to pull task‑relevant audio context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jun 17

A Closer Look at Failure Modes in Temporal Understanding of Large Audio-Language Models

arXiv:2606. 17417v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) achieve strong performance on a variety of audio understanding tasks but continue to struggle with temporal reasoning, a fundamental capability central to human auditory perception.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Sarah Wiegreffe, Dinesh Manocha, Ramani Duraiswami
arXiv Machine Learning
Sep 14

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

The paper introduces STAG, a post‑hoc framework that provides token‑level spectro‑temporal grounding for captions produced by audio‑based multimodal large language models (MLLMs). STAG estimates temporal support for each token via vocabulary projections of encoded audio, measures frequency‑band relevance through controlled spectral occlusion, and fuses these signals into a spectro‑temporal relevance map. Evaluations across ten explanation methods and four grounding benchmarks show that STAG achieves superior event‑localization performance on every dataset, and counterfactual deletion experiments confirm that removing the identified evidence selectively reduces model confidence and often eliminates the corresponding event from regenerated captions.

By Lucia Cascone, Valeria Fraenza, Michele Nappi, Fabio Narducci, Benedetto Simone