arXiv Machine Learning

Learning in Structured Stackelberg Games

arXiv Machine Learning
Aug 17

What preferences can - and cannot - predict in multi-agent online learning

arXiv:2608. 13810v1 Announce Type: cross Abstract: We examine the interplay between ordinal, preference-based solution concepts in games and the long-run behavior of game dynamics, asking in particular to what extent the combinatorial data of a game -- its preference graph -- determine the outcomes of no-regret learning dynamics -- such as follow-the-regularized-leader (FTRL).

By Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos
arXiv AI
Jul 10

Provably Optimal Learning Algorithms for Assistance Games

arXiv:2607. 08012v1 Announce Type: cross Abstract: This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact over $T$ timesteps to optimize a common reward function.

By Nivasini Ananthakrishnan, Mark Bedaywi, Michael I. Jordan, Stuart Russell, Nika Haghtalab
arXiv Machine Learning
Aug 19

Policy Optimization and Statistical Inference for Online Contextual Matrix Games

The paper introduces online contextual matrix games, a framework that merges contextual bandits with multi‑player online games to handle dynamic contexts and strategic interactions. It presents OnGameLearn, an algorithm that balances exploration and exploitation across actions and contexts, providing statistical guarantees such as tail bounds, Nash equilibrium convergence, asymptotic normality, and sublinear regret. The authors also define a policy value for matrix games and propose a doubly robust, √T‑consistent estimator, demonstrating effectiveness through simulations and a hotel pricing case study.

By Liner Xiang, Yixin Wang, Hengrui Cai
arXiv AI
3d ago

On the Complexity of Preference-Based Bandits

The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.

By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)
arXiv Machine Learning
Jul 27

Eluder dimension: localise it!

arXiv:2601. 09825v3 Announce Type: replace Abstract: We establish a lower bound on the eluder dimension of generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret bounds.

By Alireza Bakhtiari, Alex Ayoub, Samuel Robertson, David Janz, Csaba Szepesv\'ari