Beyond Supra-Competitive Outcomes: Collusive Behaviour in Deep Reinforcement Learning for Optimal Execution Games
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper introduces CURB, a reward‑shaping framework that penalizes the total variation distance between an agent’s action distributions under cooperation and defection histories, thereby preventing collusive equilibria in repeated games. By linking empirical Q‑learning collusion to Simple Penal Codes, the authors prove that any non‑trivial SPC can be neutralized, and demonstrate CURB’s effectiveness in both tabular and deep Q‑learning settings for Bertrand and Cournot competition.
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.
arXiv:2608. 09389v1 Announce Type: cross Abstract: This note aims to serve as an entry point to the literature on learning in games, a topic with significant theoretical appeal and a wide range of applications -- from machine learning and data science to economics and beyond.
arXiv:2603. 05789v5 Announce Type: replace-cross Abstract: Repeated multi-agent interactions require evaluation metrics that capture not only payoff distributions but also their temporal organization.
arXiv:2606. 05363v1 Announce Type: cross Abstract: On a platform with many sellers, should a pricing algorithm explicitly model competitors' prices when learning demand?