arXiv AI

High Volatility and Action Bias Distinguish LLMs from Humans in Group Coordination

arXiv Machine Learning
Jun 17

Tacit Coordination of Large Language Models

arXiv:2601. 22184v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in multi-agent settings that require coordination without communication, from human-AI interaction to safety-critical scenarios.

By Ido Aharon, Emanuele La Malfa, Michael Wooldridge, Sarit Kraus
arXiv AI
Jun 17

Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models

arXiv:2504. 03991v2 Announce Type: replace-cross Abstract: Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-making.

By Siddharth Srikanth, Varun Bhatt, Boshen Zhang, Werner Hager, Charles Michael Lewis, Katia P. Sycara, Aaquib Tabrez, Stefanos Nikolaidis
arXiv AI
Aug 28

Assessing mentalization in humans and large language models

The study evaluates mentalization—the capacity to infer others’ beliefs and intentions—in large language models (LLMs) using two economic games and cognitive computational modeling. Researchers tested 2,099 LLM agents from four model families (DeepSeek, GPT‑4.1, GPT‑5, Gemini 2.0 Flash) against opponents of varying sophistication, comparing their performance to 251 human participants. Results show that LLMs exhibit distinct mentalizing behaviors that vary by model provider and size, with strategic prompting generally enhancing performance; notably, GPT‑5 agents adapt their recursive reasoning depth to match opponent sophistication, outperforming humans in one task.

By Aamir Sohail, Xintong Zhong, Arkady Konovalov, Patricia L. Lockwood, Lei Zhang
arXiv AI
Sep 15

From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration

The paper compares human group discussions with large language model (LLM) deliberation traces on various reasoning tasks, finding that both humans and LLMs exhibit an assembly bonus asymmetry where discussion benefits the average member more than the best initial member. While LLM groups mirror some outcome-level patterns of human deliberation, they differ in process-level behaviors: they tend to follow majorities, surface less unique information, and converge earlier. Interventions inspired by human group‑decision research yield modest outcome improvements but do not eliminate coordination bottlenecks.

By Ala N. Tak, Teruhisa Misu, Kumar Akash, Zhaobo K. Zheng, Kevin H. Joo, Jonathan Gratch
arXiv Machine Learning
Aug 27

Skill Issue: Are Skills Language-Invariant in LLMs?

The paper investigates whether large language models (LLMs) exhibit language‑specific skill differences by having two identical model instances play a text‑based game in different languages. Using a multilingual extension of TextArena, the authors evaluate three open‑weight models across eight languages and six games, finding that the same model can show markedly different performance—varying win–loss margins, invalid actions, and strategic choices—depending on the language interface. Analyses pinpoint language‑specific failures in spatial reasoning, card‑conditioned decisions, and optimal move selection, and demonstrate that adjusting the intermediate reasoning language can recover much of the lost performance.

By Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen