Attributing Emergence in Million-Agent Systems
arXiv:2605. 11404v2 Announce Type: replace Abstract: Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents.
arXiv:2606. 02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence.
arXiv:2605. 11404v2 Announce Type: replace Abstract: Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents.
arXiv:2606. 12502v1 Announce Type: cross Abstract: We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same category as information.
arXiv:2607. 15053v1 Announce Type: cross Abstract: The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!
The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."
The paper introduces Steiner’s taxonomy of group tasks to study how multi‑agent large language model (LLM) teams scale on disjunctive versus compensatory tasks. By modeling agents as conditionally independent given the item, it shows that plurality voting converges to the modal answer while averaging converges to the item‑level bias. Experiments with 13 open‑weight models and up to 30 agents reveal that disjunctive tasks benefit from larger teams, whereas compensatory tasks like Fermi estimation see little improvement, highlighting that task structure and aggregation method fundamentally determine team scaling.
arXiv:2608. 11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs.
arXiv:2605. 29874v2 Announce Type: replace-cross Abstract: Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape equilibrium behaviour in competitive multi-agent settings?
arXiv:2607. 18310v1 Announce Type: cross Abstract: Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent.
arXiv:2608.23541v1 Announce Type: cross Abstract: Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops (Chen et al., 2025), and mixture...
arXiv:2608.22152v1 Announce Type: new Abstract: Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than ac...
arXiv:2604.27167v3 Announce Type: replace-cross Abstract: On the named Prisoner's Dilemma under direct prompting, three larger instruction-tuned models, Llama-3-70B, Qwen2.5-32B, and Qwen2.5-72B, loc...
The study investigates how limited reading capacity and claim wording influence consensus outcomes in language‑model networks. By modeling message capacity as the number of messages an agent reads, the authors show that when agents read fewer than about 6.4 messages on average, a wrong consensus becomes unreachable. However, the wording of a claim—its inherent threshold—can override this effect, leading to incorrect consensus even when most agents start correct.