arXiv AI By Bla\v{z} Bertalani\v{c}, Carolina Fortuna

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

Read the original on arXiv AI →

arXiv:2606. 02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
arXiv AI
6d ago

Multi-agent Scaling Across Disjunctive and Compensatory Tasks

The paper introduces Steiner’s taxonomy of group tasks to study how multi‑agent large language model (LLM) teams scale on disjunctive versus compensatory tasks. By modeling agents as conditionally independent given the item, it shows that plurality voting converges to the modal answer while averaging converges to the item‑level bias. Experiments with 13 open‑weight models and up to 30 agents reveal that disjunctive tasks benefit from larger teams, whereas compensatory tasks like Fermi estimation see little improvement, highlighting that task structure and aggregation method fundamentally determine team scaling.

By Carolina Fortuna, Blaz Bertalanic