arXiv AI

Abstraction Agent

The paper introduces the Abstraction Agent, a zero‑shot pipeline that employs a large language model to automatically generate continuous strategic features from a natural‑language game description, score private states, and cluster them into abstraction buckets without any game‑specific evaluators or training data. The pipeline consists of four phases—feature discovery with calibration anchors, batched private‑state scoring, correlation‑based feature selection, and k‑means clustering—and achieves significant reductions in lifted‑strategy exploitability in heads‑up no‑limit Texas hold’em and outperforms scalar rank baselines in ROVER Trials. The method also transfers to other games such as four‑card Pot‑Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, demonstrating that it can uncover strategic concepts that align with recognized game theory insights.

arXiv AI
6d ago

Up and Down the Abstraction Ladder: Code-Based Skills for Language Agents

The paper introduces CodeHack, a library of code-based skills with natural-language descriptions designed to improve language agents in complex environments like NetHack. By allowing agents to invoke reusable skills instead of selecting individual actions, the study shows that skill-based agents nearly triple game progression and cut inference cost by 86% in zero‑shot settings, while still retaining the option to fall back on primitive actions. In reinforcement learning, skill-based agents learn faster, achieving a 7.2× larger average gain in dungeon level within the same training budget.

By Bart{\l}omiej Cupia{\l}, Jens Tuyls, Maciej Wo{\l}czyk, Davide Paglieri, Martin Klissarov, Benjamin Eysenbach, Piotr Mi{\l}o\'s, Karthik R. Narasimhan
arXiv AI
Aug 5

Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

arXiv:2608. 03420v1 Announce Type: new Abstract: Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood.

By Jakub Rada (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague), Viliam Lis\'y (AI Center, Department of Computer Science, Faculty of Electrical Engineering, Czech Technical University in Prague)
arXiv AI
Jun 2

MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution

arXiv:2606. 00017v1 Announce Type: new Abstract: Training language model agents for multi-agent strategic interaction presents a core difficulty: the quality of any action may depend on future events that never materialize, on moves that violate game rules, or on decisions made by other players.

By Aliaksei Korshuk, Alexander Buyantuev, Ilya Makarov
arXiv AI
Sep 23

Synthesizing Reactive Character Behaviors for Continuous Games via Programmatic Policy Search

The paper introduces a method for creating reactive character behaviors in continuous games as compact, human‑readable programs. It searches over a domain‑specific language that uses reactive geometric decisions and higher‑order constructs to discretize continuous behavior space, while eliminating redundant program forms through synthesis antipatterns. The approach, called agentic sketching, combines bottom‑up symbolic enumeration with top‑down guidance from a coding agent, and outperforms either technique alone on a benchmark of 14 continuous games.

By Maxim Gumin, Hsueh-Ti Derek Liu, Victor Zordan, Daniel Ritchie
arXiv AI
Sep 17

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.

By Joey Xiao, Haonan Huang
arXiv AI
Jul 9

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

arXiv:2607. 06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures.

By Kabir Moghe, Peter Chin