The paper presents an uncoupled learning algorithm, higher-order optimism with discounting (HOOD), for arbitrary N-player normal form games with up to K actions per player. HOOD achieves an individual regret bound of O(N³ log² K) uniformly over the play horizon by combining a discounted (N+1)-th order predictor with entropic regularization over a lifted strategy space. This design mitigates large oscillations in play, addressing a key challenge in prior attempts to attain constant regret in general games.
By Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos
arXiv:2609.21976v1 Announce Type: cross
Abstract: We introduce Multiplicatively Optimistic Regret Matching (MORM), an uncoupled learning rule for finite general-sum games. Under simultaneous full-inf...
By Ashkan Soleymani, Georgios Piliouras
arXiv:2609.22839v1 Announce Type: cross
Abstract: Can simple learning rules keep their regret bounded in self-play? Recent work achieves constant regret bounds through modified regularization and hig...
By Junsoo Ha
arXiv:2609.16751v2 Announce Type: replace-cross
Abstract: We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve c...
By Tung Mai
arXiv:2609. 16751v1 Announce Type: cross Abstract: We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon $T$.
By Tung Mai
We give deterministic and uncoupled learning dynamics for finite multiplayer general-sum games under full-information feedback that achieve constant individual swap regret, independent of the horizon...