arXiv Computation and Language

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

The paper introduces a belief‑shift evaluation benchmark for large language models (LLMs) using the social‑deduction game Werewolf. By annotating suspicion and accusation messages in LLM‑played games, the authors measure how a village‑side model’s beliefs change after each message, evaluating 40 open‑weight LLMs on 1,224 annotated messages. Results show that larger models better distinguish wolves from villagers, yet accusations still heavily sway beliefs, especially when the accuser is trusted, and even when the accuser is wolf‑aligned. "whyItMatters":"The study highlights that current open‑weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication, revealing limitations in their belief‑updating capabilities in complex social contexts."

arXiv AI
Sep 1

QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents

arXiv:2605.27068v2 Announce Type: replace-cross Abstract: Social deduction games have become a popular testbed for probing reasoning, deception, coordination, and belief modeling in Large Language Mo...

By Ye Yuan, Rui Song, Weien Li, Zeyu Li, Haochen Liu, Xiangyu Kong, Changjiang Han, Yonghan Yang, Zichen Zhao, Zixuan Dong, Fuyuan Lyu, Bowei He, Haolun Wu, Jikun Kang, Xue Liu
arXiv AI
Sep 21

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.

By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv AI
Aug 6

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

arXiv:2608. 04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike.

By S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath
arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi