arXiv Machine Learning

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

arXiv:2608. 08268v1 Announce Type: cross Abstract: As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics.

arXiv Machine Learning
1d ago

Oblivious Learning and Collusive Pricing

The paper investigates whether pricing algorithms on multi‑seller platforms should incorporate competitors’ prices when learning demand. It compares two strategies: informed sellers that use competitor prices in their learning models, and oblivious sellers that ignore them. The study finds that oblivious sellers must explore prices more aggressively to offset missing competitor information; when all sellers are oblivious, prices eventually converge to the competitive outcome, but insufficient exploration can create many pseudo‑equilibria. In mixed markets, informed sellers earn more, and the unique Nash equilibrium is a fully informed market where prices efficiently converge to the competitive outcome, showing that oblivious modeling does not reliably produce collusion.

By Yuhang Wu, Assaf Zeevi
arXiv Machine Learning
Sep 2

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

NashDreamer is a new model-based reinforcement learning framework designed for two-player zero-sum imperfect-information games. It introduces a centralized Multi-Agent Recurrent State-Space Model that separates environment dynamics from player strategy effects, enabling the use of any policy gradient algorithm while preserving convergence guarantees to Nash equilibria. Experiments on four benchmark games show that NashDreamer achieves significantly better sample efficiency than model-free baselines early in training, and the authors analyze its optimization landscape, noting a potential vulnerability to posterior collapse in stochastic settings.

By Tom\'a\v{s} Hole\v{c}ek, Viliam Lis\'y
arXiv AI
Sep 17

Decentralized Optimal Equilibrium Learning Over Dynamic Networks

The paper introduces a decentralized learning framework for finding socially optimal equilibria in finite normal-form games played over dynamic communication networks. Agents only observe their own payoffs, lack prior knowledge of the game, and communicate with time-varying neighbors using low-bandwidth, time-stamped tables instead of raw actions or payoff data. The proposed dynamics combine randomized semantic signals, table fusion, and temporal majority reconstruction to achieve finite-time logarithmic regret guarantees for optimal equilibrium selection under utilitarian and proportional-fair social welfare objectives, as demonstrated by simulations.

By Seref Taha Kiremitci, Muhammed O. Sayin
arXiv Machine Learning
Aug 31

Aspiration-based Perturbed Learning Automata in Games with Noisy Utility Measurements. Part A: Stochastic Stability in Non-zero-Sum Games

The paper introduces aspiration-based perturbed learning automata (APLA), a payoff‑based learning scheme that incorporates an aspiration factor to reinforce action selection in distributed multi‑player games. It presents a stochastic stability analysis of APLA in positive‑utility games with noisy observations, establishing that the infinite‑dimensional Markov chain induced by the dynamics can be reduced to a finite‑dimensional one. This work extends previous results beyond potential and coordination games to generic non‑zero‑sum games, with a second part focusing on weakly acyclic games.

By Georgios C. Chasparis
arXiv Machine Learning
Aug 19

Policy Optimization and Statistical Inference for Online Contextual Matrix Games

The paper introduces online contextual matrix games, a framework that merges contextual bandits with multi‑player online games to handle dynamic contexts and strategic interactions. It presents OnGameLearn, an algorithm that balances exploration and exploitation across actions and contexts, providing statistical guarantees such as tail bounds, Nash equilibrium convergence, asymptotic normality, and sublinear regret. The authors also define a policy value for matrix games and propose a doubly robust, √T‑consistent estimator, demonstrating effectiveness through simulations and a hotel pricing case study.

By Liner Xiang, Yixin Wang, Hengrui Cai