arXiv AI By Tengfei Shao

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Read the original on arXiv AI →

The study compares human deliberation in Wason selection tasks with large language model (LLM) agent groups that are seeded with participants’ pre-discussion beliefs. Across various scoring definitions, human consensus rates ranged from 24.0% to 57.0%, whereas LLM agents consistently achieved higher consensus, with gaps of 34–44 percentage points in two sensitivity analyses. Even when early stopping was removed or memorizable answers were eliminated, LLM groups still reached near-unanimous agreement, often on incorrect answers, indicating that simulated consensus does not reflect collective accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

The paper compares human group discussions with large language model (LLM) deliberation traces on various reasoning tasks, finding that both humans and LLMs exhibit an assembly bonus asymmetry where discussion benefits the average member more than the best initial member. While LLM groups mirror some outcome-level patterns of human deliberation, they differ in process-level behaviors: they tend to follow majorities, surface less unique information, and converge earlier. Interventions inspired by human group‑decision research yield modest outcome improvements but do not eliminate coordination bottlenecks.

By Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch
arXiv AI
Jun 2

Demystifying Multi-Agent Debate: The Role of Confidence and Diversity

arXiv:2601. 19921v2 Announce Type: replace-cross Abstract: Multi-agent debate (MAD) is widely used to improve large language model (LLM) performance through test-time scaling, yet recent work shows that vanilla MAD often underperforms simple majority vote despite higher computational cost.

By Xiaochen Zhu, Caiqi Zhang, Yizhou Chi, Tom Stafford, Nigel Collier, Andreas Vlachos
Hugging Face Trending Papers
Jul 21

MeetingToM: Evaluating Multimodal LLMs on Theory-of-Mind Reasoning in Multi-Party Meetings

Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics.

arXiv Computation and Language
Sep 3

AI agents reshape consensus formation in human groups

The study investigates how large language model (LLM) agents influence consensus formation in mixed human‑AI groups during a collaborative description game. Three regimes emerge: low agent proportions lead to human‑led consensus, intermediate proportions disrupt convergence, and high proportions produce strong, agent‑led consensus. The resulting consensus differs in semantic grounding and communicative form, with human‑led consensus being concrete and holistic, and agent‑led consensus being abstract and geometrically segmented.

By Lin Chen, Ziyi Liu, Xia Hu, Yong Li