Play Like Champions: Counterfactual Feedback Generation in Latent Space
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
arXiv:2607. 00190v1 Announce Type: cross Abstract: Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games.
The paper presents a systematic mapping of recent chess research involving humans, engines, neural and reinforcement‑learning systems, large language models (LLMs), and hybrid approaches. It identifies 84 core study families and classifies them by agent type, strategic‑reasoning stages, and evaluation dimensions, highlighting a strong focus on situation assessment, evaluation, and action selection while noting gaps in planning, explanation, metacognition, and human–AI collaboration. The study also distinguishes hybrid systems by integration timing and cautions that improved human performance in evaluations does not automatically imply human–AI synergy.
The paper introduces the "convention gap" as a metric for measuring implicit communication in cooperative AI, defined as the difference between predicted failure probability from literal messages and observed failure rates. Using the card game Hanabi, the authors analyze 101,000 play actions from human-human, AI-AI, and human-AI datasets, finding a +26.2pp gap in human pairs, a -0.7pp gap in AI pairs, and a +16.4pp gap in human-AI pairs, with the largest gaps occurring on plays with no hints. The study shows that convention compatibility, rather than raw AI-AI performance, may better predict an AI’s effectiveness with human partners.
arXiv:2608. 09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents.
arXiv:2604.02578v2 Announce Type: replace-cross Abstract: Humans exhibit remarkable abilities to coordinate in groups. As large language models (LLMs) become more capable, it remains an open question...
arXiv:2606.08081v2 Announce Type: replace-cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific expressions grou...
arXiv:2508. 13213v4 Announce Type: replace Abstract: Strategic decision-making requires balancing immediate opportunities against long-term objectives: a tension fundamental to competitive environments.
arXiv:2511. 04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences.
arXiv:2606. 26267v1 Announce Type: new Abstract: Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess.
arXiv:2607. 05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions.
arXiv:2606. 08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
arXiv:2510. 23216v4 Announce Type: replace Abstract: While several high profile video games have served as testbeds for Deep Reinforcement Learning (DRL), this technique has rarely been employed by the game industry for crafting authentic AI behaviors.