SwarmBench is a new benchmark designed to evaluate large language models (LLMs) as orchestrators of agent swarms, assessing accuracy, efficiency, cost, and process quality. The study finds significant variations in orchestration performance among current models, affecting not only final outcomes but also the quality of the orchestration process itself. To address these gaps, the authors introduce SwarmExp, a method that uses experience extraction and replay to consistently enhance LLM orchestration performance.
By Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao
A technical deep dive into the Codex agent loop, explaining how Codex CLI orchestrates models, tools, prompts, and performance using the Responses API.
arXiv:2606. 01351v1 Announce Type: new Abstract: The transition from single-turn models to Multi-Agent Systems (MAS) promises enhanced problem-solving capabilities, yet the centralized orchestration topology remains a critical point of fragility.
By Junze Zhu, Weihao Chen, Xuanwang Zhang, Zhen Wu, Xinyu Dai
arXiv:2606. 31518v1 Announce Type: new Abstract: Agentic Business Process Management has gained momentum recently.
By Stefanie Rinderle-Ma, Juergen Mangler, Johannes Loebbecke, Dominik Voigt, Nataliia Klievtsova, Matthias Ehrendorfer
arXiv:2608. 05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown.
By Yidian Chen, Yingzi Gu, Natan Vidra, Spurthi Setty, Sharon Zheng
Learn how Jason Liu uses Codex to preserve context, manage complex projects, and help work continue beyond a single prompt.
arXiv:2607. 02873v1 Announce Type: cross Abstract: Large language model agents driving security tool suites over the Model Context Protocol are increasingly common.
By Romain Gerard, Assmaa Zeghaider, Yan Guo
arXiv:2607. 25656v1 Announce Type: new Abstract: Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-agent systems (MAS).
By Zhenzhen Ren, Jiyan He, Xinpeng Zhang, Zhenxing Qian, Ke Han, Shuxin Zheng, GuoBiao Li, Xiaoqing Zhang
arXiv:2607. 02807v1 Announce Type: new Abstract: Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems.
By Yuvraj Virk, Zack Edds, Chunqiu Steven Xia, Lingming Zhang
arXiv:2511.15755v3 Announce Type: replace
Abstract: Large language models (LLMs) promise to accelerate incident response in production systems, yet single-agent approaches generate vague, unusable re...
By Philip Drammeh
arXiv:2606. 13707v1 Announce Type: new Abstract: The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent systems, highlighting the importance of agent orchestration for task decomposition and collaboration.
By Fan Zhang, Vireo Zhang, Shengju Qian, Haoxuan Li, Hao Wu, Jinyang Wu, Donghao Zhou, Zhihong Zhu, Zheng Lian, Xin Wang, Pheng-Ann Heng
arXiv:2606. 01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory, and recovery.
By Rahul Suresh Babu, Adarsh Agrawal