arXiv Computation and Language By Yekun Chai, Qiwei Peng, Haoyi Xiong

Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents

Read the original on arXiv Computation and Language →

The paper introduces XiangqiBench, an executable benchmark for evaluating large language model agents in Chinese chess. It tracks 8,568 multi‑turn trajectories from 12 frontier LLMs, revealing that metrics such as the Conversion Gap, Consistency Gap, and Simulation Gap overstate true closed‑loop success. The study shows that simply naming the correct move is insufficient; agents must reliably carry a plan through to a verified outcome against an opponent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 26

Confident at the moment of action: belief miscalibration in LLM play under hidden information

The paper investigates whether large language models (LLMs) correctly gauge their confidence when acting in a hidden‑information chess variant. In experiments where the location of a hidden royal piece is repeatedly relocated, the models’ stated probabilities about the piece’s position were almost never accurate at high confidence levels, with a calibration deficit concentrated in those high‑confidence events. Across multiple model configurations and providers, the same pattern emerged, and conventional evaluation metrics such as legality, cost, latency, and completion rate were found to be uncorrelated with belief quality, yet a model could still win the game despite poor confidence estimates.

By Bhushan Kashinath Joshi