arXiv:2606. 15141v1 Announce Type: cross Abstract: While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning.
By Siyuan Zhang, Jian Zong, Junyu Wang, Peiyuan Jiang, Jiahao Yan, Jingyu Zhang, Tianrui Wang, Xiaobao Wang, Longbiao Wang, Jianwu Dang
arXiv:2608.22310v1 Announce Type: new
Abstract: Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing methods often rely on LLMs to compress or rewrite...
By Yuanhua Lin, Yile Li, Zhiyuan Zhao, Jing Shang, Jian Sun
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv:2602. 14612v4 Announce Type: replace-cross Abstract: Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding.
By Kartik Hegde, Arvind Krishna Sridhar, Naveen Vakada, Yinyi Guo, Erik Visser
arXiv:2606. 30682v1 Announce Type: cross Abstract: Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space.
By Fengjie Lu, Chenang Jiang, Jiarui Hai, Helin Wang, Aaron Yee
arXiv:2609.22452v1 Announce Type: new
Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
By Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin
PRISM is an agentic retrieval framework that uses large language models in a structured loop to improve evidence gathering for multi‑hop question answering. It splits retrieval into three specialized agents—a Question Analyzer, a Selector focused on precision, and an Adder focused on recall—whose iterative interaction yields a compact yet comprehensive evidence set. Experiments on HotpotQA, 2WikiMultiHopQA, MuSiQue, and MultiHopRAG show that PRISM consistently outperforms strong baselines by achieving higher retrieval accuracy and filtering out distracting content.
By Md Mahadi Hasan Nahid, Davood Rafiei
VoiceTrace introduces a new benchmark, VoiceTrace-Bench, for hybrid speech retrieval that combines a textual query specifying "what" to retrieve with a reference speech specifying "who" to retrieve. The authors propose a two‑stage framework: VoiceTrace‑Emb, which learns unified audio‑text embeddings for efficient large‑scale retrieval, and VoiceTrace‑Reranker, which fine‑grains relevance by jointly examining query‑candidate pairs. Experiments show VoiceTrace outperforms existing methods on both traditional semantic speech retrieval benchmarks and the new hybrid setting.
By Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun
arXiv:2602. 18452v3 Announce Type: replace-cross Abstract: As conversational multimodal AI tools are increasingly adopted to process patient data for health assessment, robust benchmarks are needed to measure progress and expose failure modes under realistic conditions.
By Gaia A. Bertolino, Yuwei Zhang, Tong Xia, Domenico Talia, Cecilia Mascolo
arXiv:2606. 14591v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet they still struggle with complex audio reasoning.
By Hui Geng, Yi Su, Han Yin, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Zijian Gao, Hengzhu Liu, Xie Chen, Kele Xu
arXiv:2606. 04555v1 Announce Type: cross Abstract: Long-horizon conversational agents need to interact with users through evolving events, tasks, and goals.
By Yifan Simon Liu, Liam Gallagher, Faeze Moradi Kalarde, Jiazhou Liang, Armin Toroghi, Scott Sanner
arXiv:2606. 13544v1 Announce Type: cross Abstract: Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations.
By Soumyajit Mitra, Prabhat Pandey, Abhinav Jain, Shanmukha Sahith, K V Vijay Girish