arXiv AI By Luis Leal

The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact

Read the original on arXiv AI →

arXiv:2607. 17543v1 Announce Type: new Abstract: In two-player zero-sum games whose Nash equilibria form a convex set, regularized solvers such as Regularized Nash Dynamics (R-NaD) empirically select the maximum-entropy member: the information projection (I-projection) of a uniform reference onto the Nash set.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy

The paper investigates how a reference policy can be used to steer regularized self‑play toward a specific equilibrium in two‑player zero‑sum games. By anchoring the reference at a target equilibrium and refining the self‑play process, the authors achieve precise convergence to that target with very low exploitability and coordinate error. The study also explores the effects of off‑manifold references, mirror‑step sizing, and boundary saturation on selection accuracy.

By Luis Leal