OpenAI Blog

An open-source spec for orchestration: Symphony

Learn how Symphony, an open-source spec for Codex orchestration, turns issue trackers into always-on agent systems—boosting engineering output and reducing context switching.

arXiv Computation and Language
Sep 1

SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

SwarmBench is a new benchmark designed to evaluate large language models (LLMs) as orchestrators of agent swarms, assessing accuracy, efficiency, cost, and process quality. The study finds significant variations in orchestration performance among current models, affecting not only final outcomes but also the quality of the orchestration process itself. To address these gaps, the authors introduce SwarmExp, a method that uses experience extraction and replay to consistently enhance LLM orchestration performance.

By Jinshan Gao, Zhuoran Jin, Tianyi Men, Kang Liu, Jun Zhao
arXiv AI
Aug 7

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality

arXiv:2608. 05263v1 Announce Type: new Abstract: Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade began, or which routing decision caused the breakdown.

By Yidian Chen, Yingzi Gu, Natan Vidra, Spurthi Setty, Sharon Zheng
arXiv AI
Jun 15

Orchestra-o1: Omnimodal Agent Orchestration

arXiv:2606. 13707v1 Announce Type: new Abstract: The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent systems, highlighting the importance of agent orchestration for task decomposition and collaboration.

By Fan Zhang, Vireo Zhang, Shengju Qian, Haoxuan Li, Hao Wu, Jinyang Wu, Donghao Zhou, Zhihong Zhu, Zheng Lian, Xin Wang, Pheng-Ann Heng