arXiv AI

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

The study compares human deliberation in Wason selection tasks with large language model (LLM) agent groups that are seeded with participants’ pre-discussion beliefs. Across various scoring definitions, human consensus rates ranged from 24.0% to 57.0%, whereas LLM agents consistently achieved higher consensus, with gaps of 34–44 percentage points in two sensitivity analyses. Even when early stopping was removed or memorizable answers were eliminated, LLM groups still reached near-unanimous agreement, often on incorrect answers, indicating that simulated consensus does not reflect collective accuracy.

arXiv AI
Sep 15

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

The paper compares human group discussions with large language model (LLM) deliberation traces on various reasoning tasks, finding that both humans and LLMs exhibit an assembly bonus asymmetry where discussion benefits the average member more than the best initial member. While LLM groups mirror some outcome-level patterns of human deliberation, they differ in process-level behaviors: they tend to follow majorities, surface less unique information, and converge earlier. Interventions inspired by human group‑decision research yield modest outcome improvements but do not eliminate coordination bottlenecks.

By Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch
arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
Hugging Face Trending Papers
Jul 21

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics.

arXiv Computation and Language
Sep 3

AI agents reshape consensus formation in human groups

The study investigates how large language model (LLM) agents influence consensus formation in mixed human‑AI groups during a collaborative description game. Three regimes emerge: low agent proportions lead to human‑led consensus, intermediate proportions disrupt convergence, and high proportions produce strong, agent‑led consensus. The resulting consensus differs in semantic grounding and communicative form, with human‑led consensus being concrete and holistic, and agent‑led consensus being abstract and geometrically segmented.

By Lin Chen, Ziyi Liu, Xia Hu, Yong Li
arXiv AI
Sep 21

How do LLMs Compute Verbal Confidence

arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...

By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c
arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
Hugging Face Trending Papers
Jun 30

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM).