arXiv AI

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding

arXiv:2608. 16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks.

arXiv Computation and Language
Sep 23

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh is a new multi‑agent harness that eliminates a central orchestrator by letting workers self‑organize through a continuous cooperation loop. The system uses a shared workspace, message interface, and shared context to coordinate tasks, verify results, and merge progress asynchronously. Experiments on ProgramBench and pandoc show that scaling from 1 to 1,024 agents improves test‑pass rates by up to 49% relative, demonstrating that agent count is a viable scaling dimension for complex tasks.

By Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia, Furu Wei
arXiv AI
2d ago

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

The paper investigates how coordination among AI agents serving different users degrades performance compared to a single coordinating agent. Across five advanced models and 77 scenarios in four shared-resource environments—API key budgets, clinic calendars, personal assistant bookings, and merge queues—the study finds that multi‑agent teams consistently underperform, sometimes collapsing entirely, and that even with communication channels coordination overhead remains significant. The authors identify specific failure modes such as stalling, action overriding, and claim fabrication, and propose environment‑specific mitigations like team leads and procedural instructions, while releasing the MAMUBench benchmark for future research.

By Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
arXiv AI
Aug 26

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

AgentRoom introduces a real‑time collaborative editing protocol that enables concurrent coding by multiple large language model agents within a CRDT‑backed shared workspace. By providing file‑level claim, status, and broadcast tools, it allows agents to coordinate directly rather than relying on serial phase handoffs or independent sampling. Experiments with five frontier coding‑CLI models show that AgentRoom reduces task abandonment and run‑to‑run variation compared to solo or parallel‑merge approaches, highlighting the importance of coordination over mere parallelism.

By Seonglae Cho, Donghyun Lee
arXiv AI
Sep 21

Scaling Discovery through Test-Time Communication

The paper demonstrates that test‑time communication among agents can significantly outperform independent parallel attempts on complex tasks. In experiments on the ARC‑AGI‑3 benchmark, a team of $k$ communicating agents matched the success rate of $4k$ independent agents, with the advantage growing as the team size increased. The study also shows that communication enables solving tasks that no single agent can solve, and that these benefits transfer to research‑oriented problems such as polyomino packing and MNIST classifier compression, where communicating agents surpassed prior best scores.

By Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos
arXiv AI
Jul 1

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

arXiv:2606. 31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.

By Kaiwen Xiong, Haonian Ji, Shi Qiu, Zeyu Zheng, Cihang Xie, Xinyu Ye, Huaxiu Yao