arXiv AI By King Yeung Tsang, Zihao Zhao, Vishal Venkataramani, Haizhou Shi, Zixuan Ke, Semih Yavuz, Shafiq Joty, Hao Wang

Reward Modeling for Multi-Agent Orchestration

Read the original on arXiv AI →

arXiv:2606. 13598v1 Announce Type: new Abstract: Multi-Agent Systems (MAS) built on Large Language Models (LLMs) require effective orchestration to coordinate specialized agents, yet training such orchestrators is hindered by limited supervision and high computational cost.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

SwarmBench is a new benchmark designed to evaluate large language models (LLMs) as orchestrators of agent swarms, assessing accuracy, efficiency, cost, and process quality. The study finds significant variations in orchestration performance among current models, affecting not only final outcomes but also the quality of the orchestration process itself. To address these gaps, the authors introduce SwarmExp, a method that uses experience extraction and replay to consistently enhance LLM orchestration performance.

By Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao
arXiv Computation and Language
Sep 22

A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents

arXiv:2601.12538v2 Announce Type: replace-cross Abstract: Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) d...

By Tianxin Wei, Ting-Wei Li, Zhining Liu, Xuying Ning, Ze Yang, Jiaru Zou, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Dongqi Fu, Zihao Li, Mengting Ai, Duo Zhou, Wenxuan Bao, Yunzhe Li, Gaotang Li, Cheng Qian, Yu Wang, Xiangru Tang, Yin Xiao, Liri Fang, Hui Liu, Xianfeng Tang, Yuji Zhang, Chi Wang, Jiaxuan You, Heng Ji, Hanghang Tong, Jingrui He
arXiv AI
3d ago

EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment

EvoSteer introduces an online self‑evolving graph orchestration framework that continuously builds and repairs a team of tool‑using agents during execution. It employs Anchored Trajectory Balance (AnchorTB), a regression‑style loss that assigns credit to each orchestration action by comparing subtrajectories to a frozen reference, and Validated Skill Admission, which tests candidate skills before promotion. Experiments on twelve datasets demonstrate that EvoSteer outperforms existing baselines in question answering, mathematical reasoning, code generation, and interactive decision making.

By Mingda Zhang, Hanwen Zhang, Qiang Huang, Zijia Wang, Pengfei Guo, Yuchen Zhang, Jionghao Zhu, Xiaoying Tang