UniAE-MoE is a unified audio encoder that uses a Mixture-of-Experts architecture to model cross‑domain audio representations. It integrates encoder components from Qwen2‑Audio and Audio‑Flamingo 3, enhances them with SwiGLU and shared experts, and applies a two‑stage instruction‑tuning strategy along with task‑specific data scaling. The model achieves state‑of‑the‑art results on the XARES‑LLM benchmark (0.802) and tops the Interspeech 2026 Audio Encoder Capability Challenge, demonstrating strong generalization across speech, music, and general audio tasks.
By Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu
AdaLoop is a lightweight recurrent module that adaptively determines how many latent refinement steps are needed for audio–question pairs, allowing deeper reasoning only when necessary. It shares a transformer block that iterates over the audio representation guided by the question, with a learned halting mechanism that exits the loop once the representation is ready. Adding fewer than 3 % of the base model’s parameters, AdaLoop improves average accuracy by 2.9 to 3.8 points across three distinct models, especially on perception-heavy subtasks.
By Lee Seung-woo, Bowen Qi
The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.
By Yotaro Kubo, Qi Sun, Yujin Tang
arXiv:2603. 09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored.
By Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee
arXiv:2606. 14591v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet they still struggle with complex audio reasoning.
By Hui Geng, Yi Su, Han Yin, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Zijian Gao, Hengzhu Liu, Xie Chen, Kele Xu
arXiv:2609.22771v1 Announce Type: cross
Abstract: Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as s...
By Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin