arXiv AI By Siu Tung Wong (Institute of Finance and Technology, University College London), Carlo Campajola (Institute of Finance and Technology, University College London, UZH Blockchain Center)

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

Read the original on arXiv AI →

The paper investigates whether a proximal policy optimisation (PPO) agent can learn to act as a broker in a continuous‑time broker‑trader game that has an analytically solved optimal strategy. By discretising the continuous‑time reward and validating the implementation, the authors show that a PPO agent with a feed‑forward neural network can approximate the reference action when there is no uninformed order flow, but struggles under stochastic uninformed flow, with critics failing to rank actions reliably. The study further demonstrates that a causal certainty‑equivalent controller based on observable history outperforms PPO under partial information, and that freezing the analytical policy and fine‑tuning PPO after a change in execution cost yields a measurable improvement. whyItMatters:"The analytical solution serves both as a diagnostic benchmark for RL performance and as a practical starting policy that can be adapted to changing market conditions, illustrating how theoretical finance models can guide and improve reinforcement learning in complex financial control tasks."

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 9

Can Reinforcement Learning Efficiently Discover Price Manipulation?

arXiv:2607. 06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates.

By Ioanna-Yvonni Tsaknaki, Andrea Macr\`i, Fabrizio Lillo
arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv Machine Learning
Jul 14

Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.

By Wen-Ting Wang
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv Machine Learning
Sep 3

Reinforcement learning to choose optimizers

The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.

By Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa