arXiv AI By Yufeng Wang

Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents

Read the original on arXiv AI →

arXiv:2606. 00476v1 Announce Type: new Abstract: Do LLM agents act on the reasoning they state?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

Why Do LLMs Struggle in Strategic Play? Broken Links Between Observations, Beliefs, and Actions

The paper investigates why large language models (LLMs) struggle in strategic decision-making under incomplete information. It identifies two key gaps: an observation‑belief gap where LLMs’ internal representations of game states are accurate but brittle, and a belief‑action gap where converting these internal beliefs into actions is weak, leading to suboptimal payoffs. Experiments with Llama 3.1, Qwen3, and gpt‑oss confirm that acting optimally on decoded beliefs would improve outcomes in most games, highlighting a bottleneck in belief‑to‑action conversion.

By Jan Sobotka, Mustafa O. Karabag, Ufuk Topcu
arXiv Computation and Language
1d ago

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

The paper introduces XiangqiBench, an executable benchmark for evaluating large language model agents in Chinese chess. It tracks 8,568 multi‑turn trajectories from 12 frontier LLMs, revealing that metrics such as the Conversion Gap, Consistency Gap, and Simulation Gap overstate true closed‑loop success. The study shows that simply naming the correct move is insufficient; agents must reliably carry a plan through to a verified outcome against an opponent.

By Yekun Chai, Qiwei Peng, Haoyi Xiong
arXiv AI
Jun 16

Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP

arXiv:2606. 16014v1 Announce Type: cross Abstract: Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-authored narrative with deeply-simulated worlds -- most acute in sandbox and open-world settings -- has been prohibitively expensive.

By Yuhang Huang, Chenmiao Li, Chaowei Fang