arXiv Machine Learning

Distributed GNEP Algorithms without Multiplier Sharing and Applications to Multi-Robot Coordination and Contextual Bandit-Based Active Learning

arXiv:2606. 00759v1 Announce Type: new Abstract: Recent advances in artificial intelligence have expanded the focus from classical optimization to include equilibrium analysis in noncooperative games.

arXiv AI
Sep 1

Fully Distributed GNE Algorithms for Multi-Robot Placement without Consensus on Multipliers

The paper introduces a fully distributed continuous‑time algorithm for solving Generalized Nash Equilibrium Problems (GNEPs) with shared linear equality constraints. Unlike existing methods that require exchanging Lagrange multipliers, this approach converges to any GNE without multiplier communication, thereby reducing communication overhead and enhancing privacy. Discrete‑time variants are also presented and the method is demonstrated on a multi‑robot placement task.

By Shao-An Yin, Mingyi Hong, Nicola Elia
arXiv Machine Learning
Aug 18

Sequential Batch Learning in Finite-Action Linear Contextual Bandits

arXiv:2004. 06321v2 Announce Type: replace Abstract: We study the sequential batch learning problem in linear contextual bandits with finite action sets, where the decision maker is constrained to split incoming individuals into (at most) a fixed number of batches and can only observe outcomes for the individuals within a batch at the batch's end.

By Yanjun Han, Zhengqing Zhou, Zihao Hu, Jose Blanchet, Peter W. Glynn, Yinyu Ye, Zhengyuan Zhou
arXiv AI
Jul 29

Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling

arXiv:2607. 25835v1 Announce Type: new Abstract: Distributed constraint optimization problems (DCOPs) provide a popular framework for distributed decision making under limited communication, but many real-world instances are too large to solve monolithically.

By Itai Zilberstein, Pranav Rajbhandari, Steve Chien, Tuomas Sandholm
arXiv Machine Learning
Aug 19

Policy Optimization and Statistical Inference for Online Contextual Matrix Games

The paper introduces online contextual matrix games, a framework that merges contextual bandits with multi‑player online games to handle dynamic contexts and strategic interactions. It presents OnGameLearn, an algorithm that balances exploration and exploitation across actions and contexts, providing statistical guarantees such as tail bounds, Nash equilibrium convergence, asymptotic normality, and sublinear regret. The authors also define a policy value for matrix games and propose a doubly robust, √T‑consistent estimator, demonstrating effectiveness through simulations and a hotel pricing case study.

By Liner Xiang, Yixin Wang, Hengrui Cai
arXiv Machine Learning
Jun 5

Multi-Agent Lipschitz Bandits

arXiv:2602. 16965v2 Announce Type: replace Abstract: We study the decentralized multi-player stochastic bandit problem over a continuous, Lipschitz-structured action space where hard collisions yield zero reward.

By Sourav Chakraborty, Amit Kiran Rege, Claire Monteleoni, Lijun Chen