arXiv Machine Learning By Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos

Constant regret in general games via higher-order optimism

Read the original on arXiv Machine Learning →

The paper presents an uncoupled learning algorithm, higher-order optimism with discounting (HOOD), for arbitrary N-player normal form games with up to K actions per player. HOOD achieves an individual regret bound of O(N³ log² K) uniformly over the play horizon by combining a discounted (N+1)-th order predictor with entropic regularization over a lifted strategy space. This design mitigates large oscillations in play, addressing a key challenge in prior attempts to attain constant regret in general games.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 1

Constant Individual Regret in General Games

arXiv:2608. 31166v1 Announce Type: new Abstract: Uncoupled no-regret dynamics provide a decentralized route to equilibrium, but prior guarantees for individual regret retain a polylogarithmic dependence on the horizon.

By Mingyang Liu, Gabriele Farina, Asuman Ozdaglar