arXiv:2608. 09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents.
By Keyu He, Xuhui Zhou, Maarten Sap
arXiv:2607. 10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did.
By Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang
The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv:2608. 01425v1 Announce Type: cross Abstract: Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward.
By Yi Mao, Andrew Perrault
arXiv:2609.35928v1 Announce Type: cross
Abstract: Multi-agent LLM systems increasingly mix models from several providers, yet exposing each agent's underlying model identity to its peers significantl...
By Xavier Del Giudice, Alessio Palma, Matteo Migliarini, Fabio Galasso, Indro Spinelli
The paper introduces a belief‑shift evaluation benchmark for large language models (LLMs) using the social‑deduction game Werewolf. By annotating suspicion and accusation messages in LLM‑played games, the authors measure how a village‑side model’s beliefs change after each message, evaluating 40 open‑weight LLMs on 1,224 annotated messages. Results show that larger models better distinguish wolves from villagers, yet accusations still heavily sway beliefs, especially when the accuser is trusted, and even when the accuser is wolf‑aligned.
"whyItMatters":"The study highlights that current open‑weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication, revealing limitations in their belief‑updating capabilities in complex social contexts."
By Yu-Yu Yang, Ti-Rong Wu, Hung Guei, Hsing-Yu Chen, I-Chen Wu