arXiv AI

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

The paper investigates whether a proximal policy optimisation (PPO) agent can learn to act as a broker in a continuous‑time broker‑trader game that has an analytically solved optimal strategy. By discretising the continuous‑time reward and validating the implementation, the authors show that a PPO agent with a feed‑forward neural network can approximate the reference action when there is no uninformed order flow, but struggles under stochastic uninformed flow, with critics failing to rank actions reliably. The study further demonstrates that a causal certainty‑equivalent controller based on observable history outperforms PPO under partial information, and that freezing the analytical policy and fine‑tuning PPO after a change in execution cost yields a measurable improvement. whyItMatters:"The analytical solution serves both as a diagnostic benchmark for RL performance and as a practical starting policy that can be adapted to changing market conditions, illustrating how theoretical finance models can guide and improve reinforcement learning in complex financial control tasks."

arXiv AI
Jul 9

Can Reinforcement Learning Efficiently Discover Price Manipulation?

arXiv:2607. 06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates.

By Ioanna-Yvonni Tsaknaki, Andrea Macr\`i, Fabrizio Lillo
arXiv AI
Sep 1

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.

By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv Machine Learning
Jul 14

Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.

By Wen-Ting Wang
arXiv Machine Learning
Aug 28

Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning

The paper investigates the problem of sharing a single critic across multiple parallel environments in reinforcement learning. It shows that when environments assign different expected returns to the same state, a shared critic must reconcile conflicting value targets, which can distort advantage estimates and misguide policy updates. The authors propose a simple fix—providing the critic with the environment index—demonstrating through bandit models and experiments on CartPole, MuJoCo, BipedalWalker, and 16 Procgen games that this conditional critic stabilizes learning and boosts returns, achieving a 40.8% improvement in aggregate normalized return on unseen levels.

By Zhenya Liu, Yang Meng, Zhuokai Zhao, Xuefeng Liu, Yuxin Chen
arXiv Machine Learning
Sep 3

Reinforcement learning to choose optimizers

The paper introduces a reinforcement learning framework that selects among a portfolio of gradient‑based and derivative‑free optimizers during a run. At each decision point a recurrent policy reads the current run state and chooses both the next optimizer and its usage duration, passing the best solution and step size forward. The method is trained with a decoupled actor‑critic using the same runtime distribution metric as evaluation, and on unseen problems it outperforms all individual portfolio optimizers except at the smallest budgets, remaining robust to distribution shift.

By Martin van der Schelling, Deepesh Toshniwal, Miguel A. Bessa
arXiv AI
4d ago

PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading

PPO‑HRAP introduces a hybrid regime‑aware policy that blends Proximal Policy Optimization with a volatility‑conditioned regime prior to balance upside participation and drawdown control in trading. The agent uses market and portfolio features, rewards that combine log return, VIX‑conditioned drawdown penalty, exposure deviation, and turnover cost, and outputs a blended action between the PPO actor and the regime‑derived target exposure. In backtests on SPY (2020‑2022) it achieved a 27.62% total return, 8.48% annualized return, and reduced maximum drawdown from 34.10% to 18.47%, while maintaining stable performance across multiple seeds and ranking first on total return and Sharpe ratio in single‑run cross‑asset tests on QQQ and DIA.

By Duong Hien Chi Kien, Thanh Trung Huynh
Hugging Face Trending Papers
Aug 20

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks.

arXiv Machine Learning
Sep 2

Accelerating Reinforcement Learning via MPC Solver-Gradient Guidance for Weights-varying MPC

The paper introduces Solver-Gradient Guided Reinforcement Learning (SG‑RL), a method that augments standard RL with bounded gradients from a differentiable MPC solver to adapt cost‑function weights online. SG‑RL integrates solver‑gradient guidance into PPO through actor‑update scaling, policy loss, advantage estimation, and value‑function learning, achieving comparable or superior closed‑loop performance while requiring up to 70.6% fewer samples. Experiments on two autonomous racing platforms with intentional model mismatch demonstrate that SG‑RL outperforms both RL and gradient‑based policy learning baselines and generalizes zero‑shot to unseen environments.

By Baha Zarrouki, Arslan Thobani, Jasper Hoffmann, Mattia Piccinini, Rudolf Reiter, Felix Jahncke, S\'ebastien Gros, Davide Scaramuzza, Johannes Betz