Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System
arXiv:2606. 26267v1 Announce Type: new Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess.
arXiv:2508. 13213v4 Announce Type: replace Abstract: Strategic decision-making requires balancing immediate opportunities against long-term objectives: a tension fundamental to competitive environments.
arXiv:2606. 26267v1 Announce Type: new Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess.
The paper presents a systematic mapping of recent chess research involving humans, engines, neural and reinforcement‑learning systems, large language models (LLMs), and hybrid approaches. It identifies 84 core study families and classifies them by agent type, strategic‑reasoning stages, and evaluation dimensions, highlighting a strong focus on situation assessment, evaluation, and action selection while noting gaps in planning, explanation, metacognition, and human–AI collaboration. The study also distinguishes hybrid systems by integration timing and cautions that improved human performance in evaluations does not automatically imply human–AI synergy.
arXiv:2510. 11503v2 Announce Type: replace-cross Abstract: Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on expert-level or even super-human play.
arXiv:2609.35835v1 Announce Type: cross Abstract: With the sheer constant advancements raining down in the field of Artificial Intelligence, one particular possibility that may cross our mind is whet...
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
arXiv:2608. 07490v1 Announce Type: cross Abstract: Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction.
The paper examines the performance of evolutionary transfer learning and TD(lambda) in the three‑dimensional chess game Dragonchess. By re‑implementing the engine in C++ to accelerate play, the authors ran 10,000 games with statistical confidence, showing both adaptive methods outperform all other agents in a round‑robin tournament. The results indicate no significant performance difference between the evolved and learned evaluation functions, demonstrating the effectiveness of adaptive techniques in complex, novel game domains.
arXiv:2604.07733v2 Announce Type: replace Abstract: Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provid...
arXiv:2604.02578v2 Announce Type: replace-cross Abstract: Humans exhibit remarkable abilities to coordinate in groups. As large language models (LLMs) become more capable, it remains an open question...
arXiv:2505. 16388v2 Announce Type: replace Abstract: The serious games between humans and AI have only just begun.
The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.
arXiv:2604. 02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI.