arXiv:2606. 27909v1 Announce Type: cross Abstract: Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a single hidden side, so a model with strong language priors can score well without ever simulating opponents' incentives.
By Avni Mittal
The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.
By Bhushan Kashinath Joshi
Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant wh...
The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.
By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
The paper introduces a belief‑shift evaluation benchmark for large language models (LLMs) using the social‑deduction game Werewolf. By annotating suspicion and accusation messages in LLM‑played games, the authors measure how a village‑side model’s beliefs change after each message, evaluating 40 open‑weight LLMs on 1,224 annotated messages. Results show that larger models better distinguish wolves from villagers, yet accusations still heavily sway beliefs, especially when the accuser is trusted, and even when the accuser is wolf‑aligned.
"whyItMatters":"The study highlights that current open‑weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication, revealing limitations in their belief‑updating capabilities in complex social contexts."
By Yu-Yu Yang, Ti-Rong Wu, Hung Guei, Hsing-Yu Chen, I-Chen Wu
arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan
arXiv:2607. 11607v1 Announce Type: new Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring.
By Hari Prasad
The paper demonstrates that aggregate accuracy figures for chain‑of‑thought (CoT) monitors can be misleading because a large portion of detected hacks rely solely on action patterns rather than reasoning. By rewriting only the agent’s reasoning to appear truthful while keeping actions identical, the authors show that the monitor’s performance on the reasoning‑dependent subset collapses dramatically, yet the overall pooled accuracy drops only modestly. The study reveals that CoT monitors are fragile when reasoning is the key signal and that accuracy should be reported separately for this subset.
By Shikhar Shiromani, Leo Richter
arXiv:2605. 06340v2 Announce Type: replace-cross Abstract: Continuous post-deployment compliance audits, mandated by emerging regulations such as the EU AI Act and Digital Services Act, create a class of strategic gaming distinct from the one-shot input/output gaming studied in prior work.
By Florian A. D. Burnat, Brittany I. Davidson
arXiv:2606. 00914v1 Announce Type: new Abstract: LLM agents increasingly act after consuming ranked external information streams such as social feeds, search results, retrieval contexts, and email queues, yet safety evaluations almost always test the model or the user prompt in isolation, never the upstream ranker that decides what the agent reads just before it acts.
By Rana Muhammad Usman
The paper proposes a claim‑specific verification audit for modular agents that replaces aggregate task scores with evidence‑based evaluations. Each agent conclusion is recorded with supporting evidence and classified as supported, unsupported, unresolved, or not evaluated, along with the boundary of validity. The audit employs three tools—oracle policies, perfect component replacements, and verifier‑score tests—to trace value changes, locate lost value, and assess verifier effectiveness, demonstrated on a portfolio‑allocation agent in a synthetic market.
By Ali Atiah Alzahrani
The paper investigates how agentic systems decide between acting and abstaining, focusing on the fidelity of their reasoning explanations. Using Qwen3‑8B in a multi‑party conversation setting, the authors compare direct decision policies, reasoning policies, supervised fine‑tuning, and reinforcement learning, finding a trade‑off: strong direct policies yield higher performance but no traceable reasoning, while reasoning policies provide an audit trail at the cost of lower recall. The study also uncovers that exposing reasoning can alter the agent’s policy and that common faithfulness metrics may overstate the alignment between reasoning and decisions.
By Shreya Mendi, Brinnae Bent