arXiv Machine Learning By Dantong Chu, Xuefeng Gao, Yufei Zhang

Learning under Opponent Unawareness in Linear-Quadratic Stochastic Games

Read the original on arXiv Machine Learning →

arXiv:2608. 08268v1 Announce Type: cross Abstract: As firms increasingly deploy machine learning for strategic decision-making, understanding algorithmic interactions has become central to operations research and economics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Oblivious Learning and Collusive Pricing

The paper investigates whether pricing algorithms on multi‑seller platforms should incorporate competitors’ prices when learning demand. It compares two strategies: informed sellers that use competitor prices in their learning models, and oblivious sellers that ignore them. The study finds that oblivious sellers must explore prices more aggressively to offset missing competitor information; when all sellers are oblivious, prices eventually converge to the competitive outcome, but insufficient exploration can create many pseudo‑equilibria. In mixed markets, informed sellers earn more, and the unique Nash equilibrium is a fully informed market where prices efficiently converge to the competitive outcome, showing that oblivious modeling does not reliably produce collusion.

By Yuhang Wu, Assaf Zeevi
arXiv Machine Learning
Sep 2

NashDreamer: Model-Based Reinforcement Learning for Zero-Sum Imperfect-Information Games

NashDreamer is a new model-based reinforcement learning framework designed for two-player zero-sum imperfect-information games. It introduces a centralized Multi-Agent Recurrent State-Space Model that separates environment dynamics from player strategy effects, enabling the use of any policy gradient algorithm while preserving convergence guarantees to Nash equilibria. Experiments on four benchmark games show that NashDreamer achieves significantly better sample efficiency than model-free baselines early in training, and the authors analyze its optimization landscape, noting a potential vulnerability to posterior collapse in stochastic settings.

By Tom\'a\v{s} Hole\v{c}ek, Viliam Lis\'y