arXiv AI

Unsound Search with Policy and Value Networks in Legends of Code and Magic

arXiv AI
Sep 18

Efficient Nash Equilibrium Computation for Cybersecurity Games

The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.

By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
arXiv AI
Sep 1

PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC

PokaiTrainer is a competitive Pokémon VGC agent that scales belief‑state search to handle simultaneous, large joint action spaces and stochastic outcomes. The system uses PokaiEngine, a Rust battle engine that efficiently enumerates joint action outcomes with high accuracy, and adapts Student of Games to solve each decision as a Bayesian matrix game under a compute budget. In live Showdown play, the agent achieved a 59% win rate against a human field averaging ~1320 Elo, reaching an Elo band of 1350‑1400 and briefly entering the top 500 of the format.

By Max Yu
arXiv AI
Sep 17

Compiled Agency: Frontier General-Purpose Coding Agents Build Winning Game Players from Bare Interaction - from Flappy Bird to StarCraft II and Civilization

The paper introduces Gauntlet, a framework that lets large language models autonomously build game-playing agents from a bare contract—just a game description, raw observation/action interface, and an empty policy file. In a single session, the model experiments with the game, compiles a standalone controller, and the resulting program is evaluated on held‑out instances without further model calls. The authors demonstrate that these compiled agents can win full‑scale games such as StarCraft II and Civilization, marking the first time a language‑agent system has achieved standalone victory in such complex titles.

By Joey Xiao, Haonan Huang
arXiv Machine Learning
Jun 25

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

arXiv:2606. 26094v1 Announce Type: new Abstract: For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention.

By Babak Rahmani, Sebastian Dziadzio, Joschka Str\"uber, Sergio Hern\'andez-Guti\'errez, Matthias Bethge
Hugging Face Trending Papers
Jun 24

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

For most of scientific history, researchers studying behavior could only infer hidden mechanisms from outward actions: an inverse problem that becomes more tractable when observation is augmented by targeted intervention. We pose a computational analogue: given only behavioral traces of an agent in a game environment, can a learner reconstruct the underlying decision program as executable code, and how much does this reconstruction improve with the ability to design controlled experiments?