arXiv Computation and Language

Language Models that Play Chess and Explain Their Moves

arXiv AI
4d ago

Benchmarking Prompt Optimization of Large Language Models With Chess

arXiv:2610. 00416v1 Announce Type: new Abstract: Evaluating large language models becomes increasingly challenging as their capabilities advance: benchmarks can saturate, public test sets risk contamination, and assessing harder tasks can require expensive grading or execution infrastructure.

By Timoth\'ee Lesort, Alejandra L\'opez de Aberasturi G\'omez, Tristan Karch, Tom Veniat, Philippe Modard, Karl Tuyls, Ludovic Denoyer
arXiv AI
Sep 17

WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.

By Tyler McDonald, Ali Emami
arXiv AI
Sep 28

Self-Play Search Distillation for Large Language Model Reasoning

Self-Play Search Distillation (SPSD) is a framework that generates superhuman synthetic data by having MuZero-like networks play board games in executable environments. The search records are converted into structured reasoning chains that serve as environment‑grounded supervision for training large language models. When applied to Qwen3‑4B‑Base, SPSD improves performance on six mathematics benchmarks from 24.1 to 36.6 and raises the win rate on unseen games from 15% to 45%.

By Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni, Luca Ragazzi, Gianluca Moro, Pavlos Vougiouklis, Jeff Z. Pan, Pasquale Minervini
arXiv AI
Sep 17

What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models

The paper presents a systematic mapping of recent chess research involving humans, engines, neural and reinforcement‑learning systems, large language models (LLMs), and hybrid approaches. It identifies 84 core study families and classifies them by agent type, strategic‑reasoning stages, and evaluation dimensions, highlighting a strong focus on situation assessment, evaluation, and action selection while noting gaps in planning, explanation, metacognition, and human–AI collaboration. The study also distinguishes hybrid systems by integration timing and cautions that improved human performance in evaluations does not automatically imply human–AI synergy.

By Paolo Ciancarini, Remo Pareschi
arXiv AI
Aug 6

Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary

arXiv:2608. 04240v1 Announce Type: cross Abstract: Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true without the natural-language explanations that make such expertise educationally useful to experts and non-experts alike.

By S. Ashwin Hebbar, Peiyao Sheng, Sewoong Oh, Pramod Viswanath