Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf
Read the original on arXiv Computation and Language →The paper introduces a belief‑shift evaluation benchmark for large language models (LLMs) using the social‑deduction game Werewolf. By annotating suspicion and accusation messages in LLM‑played games, the authors measure how a village‑side model’s beliefs change after each message, evaluating 40 open‑weight LLMs on 1,224 annotated messages. Results show that larger models better distinguish wolves from villagers, yet accusations still heavily sway beliefs, especially when the accuser is trusted, and even when the accuser is wolf‑aligned. "whyItMatters":"The study highlights that current open‑weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication, revealing limitations in their belief‑updating capabilities in complex social contexts."
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.