arXiv:2605.14483v2 Announce Type: replace
Abstract: Large language models (LLMs) provide a flexible foundation for multi-agent systems, but their effectiveness and computational cost depend criticall...
By Xudong Chen, Yixin Liu, Hua Wei, Kaize Ding
SwarmBench is a new benchmark designed to evaluate large language models (LLMs) as orchestrators of agent swarms, assessing accuracy, efficiency, cost, and process quality. The study finds significant variations in orchestration performance among current models, affecting not only final outcomes but also the quality of the orchestration process itself. To address these gaps, the authors introduce SwarmExp, a method that uses experience extraction and replay to consistently enhance LLM orchestration performance.
By Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao
arXiv:2601.12538v2 Announce Type: replace-cross
Abstract: Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) d...
By Tianxin Wei, Ting-Wei Li, Zhining Liu, Xuying Ning, Ze Yang, Jiaru Zou, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Dongqi Fu, Zihao Li, Mengting Ai, Duo Zhou, Wenxuan Bao, Yunzhe Li, Gaotang Li, Cheng Qian, Yu Wang, Xiangru Tang, Yin Xiao, Liri Fang, Hui Liu, Xianfeng Tang, Yuji Zhang, Chi Wang, Jiaxuan You, Heng Ji, Hanghang Tong, Jingrui He
arXiv:2607. 23678v1 Announce Type: new Abstract: Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use.
By Mingzhou Fan, Siyuan Xu, Mingxuan Yuan
EvoSteer introduces an online self‑evolving graph orchestration framework that continuously builds and repairs a team of tool‑using agents during execution. It employs Anchored Trajectory Balance (AnchorTB), a regression‑style loss that assigns credit to each orchestration action by comparing subtrajectories to a frozen reference, and Validated Skill Admission, which tests candidate skills before promotion. Experiments on twelve datasets demonstrate that EvoSteer outperforms existing baselines in question answering, mathematical reasoning, code generation, and interactive decision making.
By Mingda Zhang, Hanwen Zhang, Qiang Huang, Zijia Wang, Pengfei Guo, Yuchen Zhang, Jionghao Zhu, Xiaoying Tang
arXiv:2601. 10560v2 Announce Type: replace-cross Abstract: Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high inference latency from multi-step execution and repeated model invocations.
By Xi Shi, Mengxin Zheng, Qian Lou