arXiv AI By Quanwei Tang, Dong Zhang, Shoushan Li, Guodong Zhou

Don't Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding

Read the original on arXiv AI →

The paper introduces the LongAudioQA dataset to support long‑form audio meeting understanding, addressing the scarcity of task‑specific question answering data. It proposes the GRGA model, which represents heterogeneous audio features as a multi‑dimensional graph and employs an agent‑planning approach for retrieval and answer generation. The work aims to overcome acoustic information loss and limited long‑term context memory in existing speech QA methods and Speech LLMs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 16

EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning

arXiv:2606. 15141v1 Announce Type: cross Abstract: While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checkable reasoning process when dealing with complex audio reasoning.

By Siyuan Zhang, Jian Zong, Junyu Wang, Peiyuan Jiang, Jiahao Yan, Jingyu Zhang, Tianrui Wang, Xiaobao Wang, Longbiao Wang, Jianwu Dang
arXiv Computation and Language
Sep 22

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

arXiv:2609.22214v1 Announce Type: new Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...

By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen