arXiv AI
Sep 18

Efficient Nash Equilibrium Computation for Cybersecurity Games

The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.

By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
arXiv AI
Sep 1

PokaiTrainer: Scaling Belief-State Search to Competitive Pok\'emon VGC

PokaiTrainer is a competitive Pokémon VGC agent that scales belief‑state search to handle simultaneous, large joint action spaces and stochastic outcomes. The system uses PokaiEngine, a Rust battle engine that efficiently enumerates joint action outcomes with high accuracy, and adapts Student of Games to solve each decision as a Bayesian matrix game under a compute budget. In live Showdown play, the agent achieved a 59% win rate against a human field averaging ~1320 Elo, reaching an Elo band of 1350‑1400 and briefly entering the top 500 of the format.

By Max Yu