SAiFE-gym is a Python module that offers simulation environments for studying trading in Constant Product Markets with Concentrated Liquidity. It decomposes the microstructure of these markets into interactive components, allowing researchers to model various economic settings. The environments are vectorized for scalability in high‑dimensional reinforcement learning workflows, and the paper demonstrates their usefulness by evaluating RL agents under uncertain market parameters.
By Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos, Leandro S\'anchez-Betancourt, Carmine Ventre
arXiv:2607. 06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates.
By Ioanna-Yvonni Tsaknaki, Andrea Macr\`i, Fabrizio Lillo
The paper investigates whether a proximal policy optimisation (PPO) agent can learn to act as a broker in a continuous‑time broker‑trader game that has an analytically solved optimal strategy. By discretising the continuous‑time reward and validating the implementation, the authors show that a PPO agent with a feed‑forward neural network can approximate the reference action when there is no uninformed order flow, but struggles under stochastic uninformed flow, with critics failing to rank actions reliably. The study further demonstrates that a causal certainty‑equivalent controller based on observable history outperforms PPO under partial information, and that freezing the analytical policy and fine‑tuning PPO after a change in execution cost yields a measurable improvement.
whyItMatters:"The analytical solution serves both as a diagnostic benchmark for RL performance and as a practical starting policy that can be adapted to changing market conditions, illustrating how theoretical finance models can guide and improve reinforcement learning in complex financial control tasks."
By Siu Tung Wong (Institute of Finance and Technology, University College London), Carlo Campajola (Institute of Finance and Technology, University College London, UZH Blockchain Center)
arXiv:2607. 02864v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful approach in financial trading, enabling agents to learn optimal strategies through direct market interaction.
By Lin Li, Li Rong Wang, Hsuan Fu, Xiuyi Fan
arXiv:2606. 04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets.
By Damian Lebied\'z, Robert \'Slepaczuk
The paper proposes an adversarial reinforcement‑learning framework for market making that incorporates Hawkes‑process driven order arrivals and trade‑induced price impact, addressing limitations of prior Poisson‑based models. An LSTM module captures temporal dependencies in recent observations to handle increased non‑stationarity, and the authors analyze equilibrium properties and introduce a robustness evaluation protocol focused on the left tail of returns. Experiments across diverse market regimes demonstrate that the method improves left‑tail performance, especially under strong Hawkes excitation and moderate price impact, without relying on a terminal inventory bias.
By Hao Yang, Zhenguo Xu
arXiv:2606. 06201v1 Announce Type: new Abstract: Pharmaceutical supply chains (PSCs) struggle with inventory management (IM) due to unpredictable demand patterns and variable lead times associated with restocking.
By Amandeep Kaur, Gyan Prakash
arXiv:2609.13825v1 Announce Type: new
Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) o...
By Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita
arXiv:2608. 02343v1 Announce Type: cross Abstract: Many operational problems are constrained sequential decision processes with large, combinatorial action spaces and interdependent feasibility constraints.
By Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner
The paper introduces reinforcement learning for Continuous-Time Jump Markov Decision Processes (CTJMDPs) with general discrete state spaces and continuous/discrete actions. It develops entropy‑regularized continuous‑time control and establishes theoretical foundations for q‑learning in this setting, providing model‑free algorithms that outperform naive discretization. Numerical tests on network dynamic pricing demonstrate the method’s ability to learn near‑optimal policies and scale to large networks.
By Huiling Meng, Ningyuan Chen, Xuefeng Gao
arXiv:2607. 10960v1 Announce Type: new Abstract: Trader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment.
By Wen-Ting Wang
arXiv:2606. 06823v1 Announce Type: cross Abstract: While deep learning has excelled in various domains, its application to sequential decision-making in finance remains challenging due to the low Signal-to-Noise Ratio (SNR) and non-stationarity of financial data.
By Yuqi Li, Siyuan Liu, Bingjun Liu