arXiv:2609.20849v1 Announce Type: new
Abstract: Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) re...
By Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli
arXiv:2607. 21550v1 Announce Type: new Abstract: While large audio-language models have achieved remarkable progress in auditory perception, they still lag behind text-based large language models in deep logical reasoning, primarily due to the scarcity of high-quality audio reasoning data.
By Dongjie Fu, Di Cao, Xize Cheng, Zihan Zhang, Wenxu Jia, Yifu Chen, Shengpeng Ji, Yu Zhang, Tao Jin
arXiv:2609.23589v1 Announce Type: cross
Abstract: Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio repr...
By Jiaheng Dong, Xiaofeng Yu, Jean Honorio, Abhirup Ghosh, Hong Jia, Ting Dang
ReasonAudio is a new benchmark designed to evaluate reasoning capabilities in text‑audio retrieval, addressing the gap left by existing semantic‑matching focused datasets. It tests four logical abilities—negation, temporal order, sound co‑occurrence, and sound duration—across five synthetic subtasks (1,000 queries over 10,000 composite clips) and one natural subtask (100 queries over 1,000 real‑world clips). Evaluation of 11 state‑of‑the‑art systems shows significant limitations, with the best model, OmniEmbed‑7B, scoring only 20.7 overall and 53.8% in a controlled setting, compared to 70.6% for its generative backbone and 95.6% for humans.
By Honglei Zhang, Yuting Chen, Chenpeng Hu, Pengfei Zhou, Siyue Zhang, Yilei Shi
arXiv:2608. 16539v1 Announce Type: cross Abstract: Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized.
By Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais, Philip J. B. Jackson, Jiankang Deng, Ismail Elezi
arXiv:2606. 30682v1 Announce Type: cross Abstract: Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space.
By Fengjie Lu, Chenang Jiang, Jiarui Hai, Helin Wang, Aaron Yee