arXiv AI

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

arXiv:2606. 02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence.

arXiv AI
4d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
arXiv AI
6d ago

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

The paper introduces Steiner’s taxonomy of group tasks to study how multi‑agent large language model (LLM) teams scale on disjunctive versus compensatory tasks. By modeling agents as conditionally independent given the item, it shows that plurality voting converges to the modal answer while averaging converges to the item‑level bias. Experiments with 13 open‑weight models and up to 30 agents reveal that disjunctive tasks benefit from larger teams, whereas compensatory tasks like Fermi estimation see little improvement, highlighting that task structure and aggregation method fundamentally determine team scaling.

By Carolina Fortuna, Blaz Bertalanic
arXiv AI
Jun 16

Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension

arXiv:2605. 29874v2 Announce Type: replace-cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibrium behaviour in competitive multi-agent settings?

By Francisco Le\'on Z\'u\~niga Bol\'ivar (Instituci\'on Universitaria Colegio Mayor del Cauca)
arXiv Computation and Language
Sep 18

Message capacity and claim wording set the transition points of collective truth-finding in language-model networks

The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.

By Makoto Fukushima