arXiv AI

When Order Matters: First-Speaker Bias and Mitigation through Personality in Sequential Multi-Agent Debate

arXiv Computation and Language
Aug 27

Reasoning or Rambling? Exploring the Effect of Thinking on Agent Persuasion

The paper investigates how explicit reasoning in Large Reasoning Models (LRMs) affects their ability to persuade and be persuaded. Experiments on objective and subjective tasks reveal a Persuasion Duality: reasoning boosts an agent’s persuasive power by about 21 percentage points while also making it less susceptible to incorrect persuasion by up to 10 percentage points. However, the study finds that persuasiveness often relies on superficial cues like response length and repetition rather than logical validity, and that persuasion can amplify or attenuate non‑linearly across multi‑hop agent chains. The authors also propose an attention‑guided prompt‑level adversarial argument detection method that improves agent robustness.

By Haodong Zhao, Jidong Li, Zhaomin Wu, Tianjie Ju, Zhuosheng Zhang, Bingsheng He, Gongshen Liu
arXiv Machine Learning
Sep 17

Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric

The paper investigates how a minority of biased agents in a multi‑agent system of large language models (LLMs) can amplify bias through textual interactions. Even a small percentage of persistently extreme agents causes significant opinion shifts among the non‑biased agents, with the effect occurring faster in the Llama 3.2 model than in a classical Friedkin‑Johnsen model. Semantic analysis shows that rhetorical consistency rises with biased exposure and that non‑biased agents adopt the biased vocabulary even when their numerical opinions change only modestly.

By Omran Berjawi, Giuseppe Fenza, Rida Khatoun
arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
arXiv AI
1d ago

AI Agents are Vulnerable to Radicalization

The study explores how large language models (LLMs) can influence each other’s beliefs by simulating conversations between a target LLM and an influencer LLM. It identifies two radicalization pathways—resonance, which amplifies pre‑existing beliefs, and persuasion, which introduces new beliefs—and finds that resonance consistently produces stronger radicalization effects. The research also shows that different influence tactics yield varying levels of radicalization and that resonance can spread to related beliefs, indicating interconnected belief structures within AI agents.

By Ozgur Can Seckin, Shalmoli Ghosh, Alessandro Flammini, Kristina Lerman, Maria Elizabeth Grabe, Filippo Menczer
arXiv AI
Sep 4

Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation

The paper introduces R$^2$-MAD, a framework that enhances multi-agent debate by giving agents an experience memory from past debates. It uses a debate-state-aware retrieval policy to adjust concept priors based on current consensus, and derives confidence weights from retrieved experiences to modulate peer influence. Experiments demonstrate consistent improvements over existing single-agent and MAD baselines.

By Xuanfa Jin, Zhijian Ma, Yongcheng Zeng, Xinyu Cui, Haifeng Zhang, Jun Wang
arXiv AI
4d ago

Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content

The study evaluates how well language‑model agents can simulate individual social media reactions by comparing predictions under different prompt conditions. Eight Serbian participants’ reactions to 68 posts were recorded, and four language models were asked to predict these reactions using prompts that varied in profile content and instruction style. The results show that prompts emphasizing attitudinal content and intuitive, immediate responses yield the highest fidelity, outperforming demographic backstories and a crowd baseline, and suggesting that such agents could act as general‑purpose simulated users.

By Ljubisa Bojic, Tijana Stanic, Joerg Matthes, Agariadne Dwinggo Samala, Bojana Dinic, Jue Wang
arXiv AI
2d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng