arXiv:2610.00619v1 Announce Type: cross
Abstract: In this paper, we extend earlier findings of supra-competitive outcomes in optimal-execution games by identifying a learned punitive mechanism that d...
By Christos Spyridon Koulouris, Carlo Campajola
arXiv:2609.13825v1 Announce Type: new
Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) o...
By Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita
arXiv:2608. 09586v1 Announce Type: new Abstract: The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values.
By Boning Li, Longbo Huang
arXiv:2606. 03736v2 Announce Type: replace-cross Abstract: We study resource-constrained dynamic pricing when the seller seeks revenue and valid inference about demand at a price fixed before the selling season.
By Ruicheng Ao, Jiashuo Jiang, David Simchi-Levi
arXiv:2510. 14642v2 Announce Type: replace-cross Abstract: In blockchain networks, the strategic ordering of transactions within blocks has emerged as a significant source of profit extraction, known as Maximal Extractable Value (MEV).
By Andrei Seoev, Leonid Gremyachikh, Anastasiia Smirnova, Yash Madhwal, Alisa Kalacheva, Dmitry Belousov, Ilia Zubov, Aleksei Smirnov, Denis Fedyanin, Vladimir Gorgadze, Yury Yanovich
The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.
By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv:2608. 09389v1 Announce Type: cross Abstract: This note aims to serve as an entry point to the literature on learning in games, a topic with significant theoretical appeal and a wide range of applications -- from machine learning and data science to economics and beyond.
By Panayotis Mertikopoulos
arXiv:2606. 05363v1 Announce Type: cross Abstract: On a platform with many sellers, should a pricing algorithm explicitly model competitors' prices when learning demand?
By Yuhang Wu, Assaf Zeevi
arXiv:2606. 02595v1 Announce Type: new Abstract: Dynamic pricing in short-term rental (STR) markets presents a distinctive challenge for online learning algorithms: pricing decisions carry significant financial risk, operators require explainability, and market feedback is sparse (one booking outcome per listed night).
By Oleg Miroshnichenko
arXiv:2608. 04832v1 Announce Type: new Abstract: Control policies optimized in simulation can perform poorly in the real system when the parameters $x$ of the simulator are estimated from limited data but the resulting parameter uncertainty is not represented inside the simulation.
By Konrad J. Mueller, Amira Akkari, Ben Wood, Lukas Gonon
arXiv:2607. 23333v1 Announce Type: cross Abstract: We revisit the regret loss framework introduced in Park et al.
By Chanwoo Park, Asuman Ozdaglar
arXiv:2606. 08791v1 Announce Type: cross Abstract: We study the problem of auditing a black-box algorithmic decision-maker from observable inputs and outputs alone.
By Irene Aldridge