arXiv:2606. 28308v1 Announce Type: cross Abstract: Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles that all share the minimax value V* yet prescribe different behaviour.
By Luis Leal
The paper investigates how a reference policy can be used to steer regularized self‑play toward a specific equilibrium in two‑player zero‑sum games. By anchoring the reference at a target equilibrium and refining the self‑play process, the authors achieve precise convergence to that target with very low exploitability and coordinate error. The study also explores the effects of off‑manifold references, mirror‑step sizing, and boundary saturation on selection accuracy.
By Luis Leal
arXiv:2608.24488v1 Announce Type: new
Abstract: Many continuous-control policies are optimized as unbounded Gaussians and then mapped into bounded actions. We show that where entropy is measured chan...
By Yiyang He, Zhichun Zhou, Ziwei Wang, Tao Xue, Haolin Fei
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2606. 16341v1 Announce Type: new Abstract: A filtered approximate-nearest-neighbor (ANN) query returns the k nearest vectors among those satisfying an attribute predicate P of selectivity s.
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2608. 04149v1 Announce Type: cross Abstract: Swap regret governs the rate at which uncoupled learning dynamics converge to correlated equilibria in multiplayer general-sum games.
By Taira Tsuchiya