When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
arXiv:2608. 16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks.
The paper demonstrates that test‑time communication among agents can significantly outperform independent parallel attempts on complex tasks. In experiments on the ARC‑AGI‑3 benchmark, a team of $k$ communicating agents matched the success rate of $4k$ independent agents, with the advantage growing as the team size increased. The study also shows that communication enables solving tasks that no single agent can solve, and that these benefits transfer to research‑oriented problems such as polyomino packing and MNIST classifier compression, where communicating agents surpassed prior best scores.
arXiv:2608. 16801v1 Announce Type: new Abstract: We study how teams of AI coding agents coordinate while solving programming tasks.
arXiv:2609.15309v1 Announce Type: new Abstract: Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to s...
DIANOIA introduces a diagnostic framework for multi‑agent large language model systems, decomposing reasoning gain into three measurable channels—coverage, fidelity, and synthesis. The protocol identifies bottleneck channels for a given task and implements a corresponding multi‑agent system with role‑diverse proposers, execution‑grounded verification, and iterative synthesis. Experiments on GSM8K, AIME‑2025, MBPP, and BFCL‑SP show that DIANOIA outperforms strong baselines, achieving significant token savings and accuracy gains while accurately pinpointing the critical channels.
arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.
arXiv:2609.22682v1 Announce Type: new Abstract: Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unkno...
The paper investigates how scaling a team of small language‑model agents affects performance across different orchestration architectures. By testing eight architectures on five short‑answer benchmarks and an executable‑code benchmark, it finds that team scaling yields large gains on arithmetic word‑problem tasks but only modest improvements on multiple‑choice and code generation tasks, with no single architecture dominating all tasks. The authors explain these patterns using a generate‑transform decomposition that separates coverage and transformation effects, showing that arithmetic tasks benefit from both coverage and critic‑guided transformation, while other tasks are limited by saturation or poor conversion.
arXiv:2608.22191v1 Announce Type: new Abstract: Software-engineering agents solve repository-level tasks through long, stochastic tool-use trajectories, and repeated attempts often find fixes missed...
arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.
arXiv:2409. 11363v2 Announce Type: replace-cross Abstract: AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research.
arXiv:2606. 31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.
arXiv:2606. 28514v1 Announce Type: new Abstract: Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents.
CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.