arXiv AI By Yuan Gao, Jiangyi Yang, Yao Zhao, Yichi Zhang

Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games

Read the original on arXiv AI →

arXiv:2607. 10814v1 Announce Type: cross Abstract: Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent decided as it did.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi
arXiv AI
Sep 21

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.

By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv Computation and Language
Sep 14

Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf

The paper introduces a belief‑shift evaluation benchmark for large language models (LLMs) using the social‑deduction game Werewolf. By annotating suspicion and accusation messages in LLM‑played games, the authors measure how a village‑side model’s beliefs change after each message, evaluating 40 open‑weight LLMs on 1,224 annotated messages. Results show that larger models better distinguish wolves from villagers, yet accusations still heavily sway beliefs, especially when the accuser is trusted, and even when the accuser is wolf‑aligned. "whyItMatters":"The study highlights that current open‑weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication, revealing limitations in their belief‑updating capabilities in complex social contexts."

By Yu-Yu Yang, Ti-Rong Wu, Hung Guei, Hsing-Yu Chen, I-Chen Wu