arXiv AI

Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy

The paper investigates how a reference policy can be used to steer regularized self‑play toward a specific equilibrium in two‑player zero‑sum games. By anchoring the reference at a target equilibrium and refining the self‑play process, the authors achieve precise convergence to that target with very low exploitability and coordinate error. The study also explores the effects of off‑manifold references, mirror‑step sizing, and boundary saturation on selection accuracy.

arXiv AI
Sep 24

ANO: Robust Policy Optimization via Bounded, Redescending Gain Fields

The paper introduces Anchored Neighborhood Optimization (ANO), a new policy‑optimization method that directly designs a smooth, bounded gain field for the probability‑ratio surrogate objective. ANO anchors the identity map at a ratio of one, peaks at a specified trust‑region boundary, and limits the influence of extreme off‑policy samples while providing a bounded, redescending pull on outliers. Empirical results show ANO consistently outperforms existing methods on Atari and MuJoCo benchmarks, and it remains robust under aggressive learning‑rate settings.

By Yiheng Zhang, Yiming Wang, Kaiyan Zhao, Zhenglin Wan, Jiayu Chen, Leong Hou U
arXiv AI
Sep 3

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

The paper argues that in multi‑turn agentic reinforcement learning, credit assignment should be viewed as a coverage problem rather than a targeting problem. It introduces verifier information density (V_d) as a structural metric, showing that terminal‑state verifiers operate in a low‑V_d regime where targeting fails. Experiments on tau^2‑bench, BFCL, and ToolACE‑2‑8B demonstrate that uniformly distributing reward across all turns outperforms sparse, targeted rewards, and that full chain coverage is necessary for optimal performance.

By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv AI
Aug 25

Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary

The paper studies a stationary decentralized Markov game where a focal agent experiences drifting rewards and dynamics due to learning peers, framing this as an agent‑centric continual reinforcement‑learning problem. It introduces the concept of an invariant core—maximal abstract patterns common to many successful trajectories—and proves a worst‑case conditioning theorem linking trajectory‑law drift to success coverage. The authors provide theoretical guarantees for survival horizon, first‑exit law, and regret, and validate their predictions with solvable models and empirical studies in continual control, cue‑MNIST, and Level‑Based Foraging.

By Dane Malenfant