arXiv Machine Learning

Multi-Agent Systems are Mixtures of Experts: Who Becomes an Influencer?

arXiv:2605. 25929v2 Announce Type: replace-cross Abstract: The effectiveness of multi-agent LLM deliberation depends not only on the agents' individual predictions, but also on how they communicate and collaborate.

arXiv Machine Learning
Sep 17

Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric

The paper investigates how a minority of biased agents in a multi‑agent system of large language models (LLMs) can amplify bias through textual interactions. Even a small percentage of persistently extreme agents causes significant opinion shifts among the non‑biased agents, with the effect occurring faster in the Llama 3.2 model than in a classical Friedkin‑Johnsen model. Semantic analysis shows that rhetorical consistency rises with biased exposure and that non‑biased agents adopt the biased vocabulary even when their numerical opinions change only modestly.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun
arXiv Computation and Language
Aug 27

Belief Cascades Drive Persuasion in LLM Agent Networks

The paper introduces a controlled testbed to study how goal‑directed persuaders shift stances in networks of large language model agents, using real‑world ego‑network topologies. Experiments across four LLM backbones, five graph structures, and 55 policy statements show that persuasion dynamics depend on topology, competition, topic, and model prior. The study finds that direct exposure predicts stance change, peer relays have measurable influence, and that post‑text analysis alone misses important movement, highlighting the need to evaluate multi‑agent persuasion through trajectory‑level processes, belief probes, exposure provenance, and action logs.

By Haoyi Qiu, Genglin Liu, Pranav Narayanan Venkit, Kung-Hsiang Huang, Saadia Gabriel, Chien-Sheng Wu, Nanyun Peng
arXiv AI
2d ago

Counting Moves, Weighing Voices: Bayesian Dialectical Argumentation for Calibrated Multi-LLM Councils under Persistent Adversaries

The paper introduces Bayesian Dialectical Argumentation (BDA), a method for aggregating answers from multiple large language models (LLMs) in a council setting. BDA treats each LLM’s typed moves—proposals, challenges, and concessions—as evidence in a classical annotator model, estimating per-agent reliability even when some agents are persistently unreliable. By weighting evidence according to these inferred reliabilities, BDA produces calibrated posterior probabilities for candidate answers and can invert unreliable agents instead of merely outvoting them, achieving superior calibration and robustness on both binary and multi-class benchmarks without extra LLM calls.

By Ionel Eduard Stan, Paolo Napoletano
arXiv AI
Sep 21

Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents

The paper introduces Bayesian Chronicle Agents (BCA), a lightweight belief layer that separates an LLM agent’s internal stance from its outward speech. Each stance is represented as a probability updated via a single Bayesian step per utterance, with a single prior‑strength parameter κ controlling stubbornness. By sweeping κ, the authors generate three controllable opinion‑dynamics regimes—consensus, persistent disagreement, and committed‑minority influence—matching Friedkin–Johnsen theory and demonstrating recoverable, auditable belief states across models.

By Hafsa Akbar, Daniel Platnick, Marjan Alirezaie, Hossein Rahnama
arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv AI
4d ago

Local Predictability and Collective Fidelity in LLM-Agent Societies

The paper investigates whether compact surrogate models can reduce the cost of simulating large language model (LLM) societies while still reproducing their collective behavior. Using 9,455 published trajectories and new opinion‑dynamics experiments, it finds that incorporating neighbor information improves individual predictions across all 16 public‑data settings and enhances pooled collective forecasts on held‑out questions, though the collective gains vary with transfer conditions. Additional tests on 24 new statements do not confirm earlier contrasting history effects, and Qwen shows benefit from history only after three observed rounds, underscoring the need for direct collective validation, explicit limits on available observations, and comparisons with simple baselines.

By Igor Itkin