arXiv AI

Concentrated Liquidity Provision: a Reinforcement Learning Perspective

arXiv:2608. 19389v1 Announce Type: cross Abstract: Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi).

arXiv AI
Sep 17

SAiFE-gym: Model-based Environments for Automated Market Making with Concentrated Liquidity

SAiFE-gym is a Python module that offers simulation environments for studying trading in Constant Product Markets with Concentrated Liquidity. It decomposes the microstructure of these markets into interactive components, allowing researchers to model various economic settings. The environments are vectorized for scalability in high‑dimensional reinforcement learning workflows, and the paper demonstrates their usefulness by evaluating RL agents under uncertain market parameters.

By Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos, Leandro S\'anchez-Betancourt, Carmine Ventre
arXiv AI
Jul 9

Can Reinforcement Learning Efficiently Discover Price Manipulation?

arXiv:2607. 06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates.

By Ioanna-Yvonni Tsaknaki, Andrea Macr\`i, Fabrizio Lillo
arXiv AI
1d ago

When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

The paper investigates whether a proximal policy optimisation (PPO) agent can learn to act as a broker in a continuous‑time broker‑trader game that has an analytically solved optimal strategy. By discretising the continuous‑time reward and validating the implementation, the authors show that a PPO agent with a feed‑forward neural network can approximate the reference action when there is no uninformed order flow, but struggles under stochastic uninformed flow, with critics failing to rank actions reliably. The study further demonstrates that a causal certainty‑equivalent controller based on observable history outperforms PPO under partial information, and that freezing the analytical policy and fine‑tuning PPO after a change in execution cost yields a measurable improvement. whyItMatters:"The analytical solution serves both as a diagnostic benchmark for RL performance and as a practical starting policy that can be adapted to changing market conditions, illustrating how theoretical finance models can guide and improve reinforcement learning in complex financial control tasks."

By Siu Tung Wong (Institute of Finance and Technology, University College London), Carlo Campajola (Institute of Finance and Technology, University College London, UZH Blockchain Center)
arXiv Machine Learning
Sep 22

Robust Market Making with Hawkes Order Flow and Price Impact via Adversarial Reinforcement Learning

The paper proposes an adversarial reinforcement‑learning framework for market making that incorporates Hawkes‑process driven order arrivals and trade‑induced price impact, addressing limitations of prior Poisson‑based models. An LSTM module captures temporal dependencies in recent observations to handle increased non‑stationarity, and the authors analyze equilibrium properties and introduce a robustness evaluation protocol focused on the left tail of returns. Experiments across diverse market regimes demonstrate that the method improves left‑tail performance, especially under strong Hawkes excitation and moderate price impact, without relying on a terminal inventory bias.

By Hao Yang, Zhenguo Xu
arXiv Machine Learning
Aug 24

Reinforcement Learning for Continuous-Time Jump Markov Decision Processes with Applications to Network Dynamic Pricing

The paper introduces reinforcement learning for Continuous-Time Jump Markov Decision Processes (CTJMDPs) with general discrete state spaces and continuous/discrete actions. It develops entropy‑regularized continuous‑time control and establishes theoretical foundations for q‑learning in this setting, providing model‑free algorithms that outperform naive discretization. Numerical tests on network dynamic pricing demonstrate the method’s ability to learn near‑optimal policies and scale to large networks.

By Huiling Meng, Ningyuan Chen, Xuefeng Gao
arXiv Machine Learning
Jul 14

Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator

arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.

By Wen-Ting Wang