arXiv:2603.02266v2 Announce Type: replace-cross
Abstract: Test-Time Scaling has shown notable efficacy in addressing complex problems through scaling inference compute. However, within Large Audio-La...
By Ruixiang Mao, Xiangnan Ma, Dan Chen, Ziming Zhu, Yuan Ge, Aokai Hao, Haishu Zhao, Yifu Huo, Qing Yang, Kaiyan Chang, Xiaoqian Liu, Chenglong Wang, Qiaozhi He, Tong Xiao, Jingbo Zhu
arXiv:2608.22236v2 Announce Type: replace-cross
Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability...
By Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang
arXiv:2606. 11400v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) excel at audio understanding but expose little about where in an audio signal they attend.
By Tsung-En Lin, Hung-Yi Lee
arXiv:2509. 22363v4 Announce Type: replace Abstract: Large Audio Language Models (LALMs) integrate audio encoders with pretrained Large Language Models to perform complex multimodal reasoning tasks.
By Pooneh Mousavi, Lovenya Jain, Mirco Ravanelli, Cem Subakan
The study applies causal tracing to large audio language models (LALMs) to uncover how they fuse acoustic and textual information. Layer‑wise analysis reveals distinct fusion strategies—progressive integration in DeSTA versus abrupt late‑stage fusion in Qwen—while token‑wise analysis identifies the final sequence token as an informational bottleneck that decisively retrieves audio content. Additionally, an attention‑like query mechanism at intermediate tokens is observed, prompting the model to pull task‑relevant audio context.
By Wei-Chih Chen, Chien-yu Huang, Hung-yi Lee
arXiv:2609.23589v1 Announce Type: cross
Abstract: Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio repr...
By Jiaheng Dong, Xiaofeng Yu, Jean Honorio, Abhirup Ghosh, Hong Jia, Ting Dang