arXiv AI

Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation

arXiv:2604. 11840v3 Announce Type: replace-cross Abstract: Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises.

arXiv AI
4d ago

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

The study evaluates multi‑agent debate (MAD) in small language models, testing whether cognitive diversity—via personas, sampling temperature, or model identity—drives performance gains. Across 23 models, five tasks, and over 5,500 runs, MAD consistently outperforms single‑model inference but, when matched for generation budget, it ties or falls behind self‑consistency sampling, with persona prompting actually reducing accuracy. The authors find that MAD’s benefits largely stem from the first answer exchange and that many reported gains are due to ensemble‑sampling effects rather than true diversity, highlighting the need for budget‑matched, contamination‑checked baselines. whyItMatters":"The findings clarify that MAD’s perceived advantages may be overestimated and that future debate mechanisms must be evaluated against rigorous, budget‑matched baselines to ensure genuine performance improvements."

By Leonardo Ferreira, Gardenia Liu, Kaden Zheng
arXiv Computation and Language
Aug 31

Benchmarking large language model agent societies against human behavioural distributions

The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.

By Raad Bin Tareaf
arXiv Machine Learning
5d ago

Reinforcement Learning of Communication in a Mesh of Small Language Models

The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.

By Mehmet Kerem Turkcan
arXiv AI
Aug 25

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

The study examines how Theory of Mind (ToM) reasoning and prosocial beliefs influence large language models (LLMs) in the ultimatum game. By initializing LLM agents with Greedy, Fair, or Selfless beliefs and applying chain‑of‑thought or varying levels of ToM reasoning, the authors ran 2,700 simulations across several models, including o3‑mini and DeepSeek‑R1 Distilled Qwen 32B. Results show that ToM‑enhanced LLMs align more closely with human decision patterns, exhibit greater consistency, and achieve better negotiation outcomes, with Llama 3.3 70B producing the most belief‑consistent reasoning. whyItMatters":"The findings clarify the importance of incorporating Theory of Mind into LLMs to improve their alignment with human norms in cooperative decision‑making tasks."

By Neemesh Yadav, Yihuai Lan, Shan Dong, Mai Hieu Hien, Palakorn Achananuparp, Jing Jiang, Ee-Peng Lim