arXiv AI By Siyuan Zhang, Jian Zong, Junyu Wang, Peiyuan Jiang, Jiahao Yan, Jingyu Zhang, Tianrui Wang, Xiaobao Wang, Longbiao Wang, Jianwu Dang

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

Read the original on arXiv AI →

arXiv:2606. 15141v1 Announce Type: cross Abstract: While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

The paper introduces the LongAudioQA dataset to support long‑form audio meeting understanding, addressing the scarcity of task‑specific question answering data. It proposes the GRGA model, which represents heterogeneous audio features as a multi‑dimensional graph and employs an agent‑planning approach for retrieval and answer generation. The work aims to overcome acoustic information loss and limited long‑term context memory in existing speech QA methods and Speech LLMs.

By Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou
arXiv AI
Aug 28

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

SpeechGym is an audio‑native environment that lets two omni‑modal models converse entirely in native audio, eliminating external ASR/TTS and API boundaries while preserving the tasks, tools, and success checks of a standard text‑based agent benchmark. By training end‑to‑end, the framework addresses perceptual failures—such as misheard arguments that cascade into failed calls—and behavioural failures, both of which are automatically labeled for free. Using per‑turn process rewards to overcome reward sparsity, agents trained in SpeechGym transfer to an independent voice benchmark, doubling task success and improving efficiency in turns and tokens.

By Jiajun Fan, Jingyuan Li, Prashanth Gurunath Shivakumar, Jia-Hong Huang, Qi Luo, M. Maruf, Ivan Bulyko, Ge Liu, Roger Ren