arXiv AI

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

The paper introduces Steiner’s taxonomy of group tasks to study how multi‑agent large language model (LLM) teams scale on disjunctive versus compensatory tasks. By modeling agents as conditionally independent given the item, it shows that plurality voting converges to the modal answer while averaging converges to the item‑level bias. Experiments with 13 open‑weight models and up to 30 agents reveal that disjunctive tasks benefit from larger teams, whereas compensatory tasks like Fermi estimation see little improvement, highlighting that task structure and aggregation method fundamentally determine team scaling.

arXiv AI
2d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
arXiv AI
Jul 1

ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents

arXiv:2606. 31174v1 Announce Type: new Abstract: Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates specialized subagents, delegates work, and orchestrates their parallel, asynchronous returns through dynamic workflows.

By Kaiwen Xiong, Haonian Ji, Shi Qiu, Zeyu Zheng, Cihang Xie, Xinyu Ye, Huaxiu Yao
arXiv AI
Jun 2

Scaling Behavior of Single LLM-Driven Multi-Agent Systems

arXiv:2606. 00655v1 Announce Type: cross Abstract: The burgeoning field of LLM-based Multi-Agent Systems (MAS) promises to tackle complex tasks through collaborative intelligence, yet fundamental questions regarding their scaling behavior and intrinsic collective dynamics remain underexplored.

By Jialing Li, Zhouhong Gu, Yin Cai, Hongwei Feng
arXiv AI
2d ago

An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration Architectures

The paper investigates how scaling a team of small language‑model agents affects performance across different orchestration architectures. By testing eight architectures on five short‑answer benchmarks and an executable‑code benchmark, it finds that team scaling yields large gains on arithmetic word‑problem tasks but only modest improvements on multiple‑choice and code generation tasks, with no single architecture dominating all tasks. The authors explain these patterns using a generate‑transform decomposition that separates coverage and transformation effects, showing that arithmetic tasks benefit from both coverage and critic‑guided transformation, while other tasks are limited by saturation or poor conversion.

By Blaz Bertalanic, Carolina Fortuna
arXiv AI
Jun 2

When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffs

arXiv:2605. 24202v2 Announce Type: replace Abstract: Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understood.

By Yifan Zeng, Yiran Wu, Yaolun Zhang, Wentian Zhao, Kun Wan, Qingyun Wu, Huazheng Wang