arXiv:2608. 19389v1 Announce Type: cross Abstract: Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi).
By Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos, Leandro S\'anchez-Betancourt, Carmine Ventre
arXiv:2607. 02864v1 Announce Type: new Abstract: Reinforcement Learning (RL) has emerged as a powerful approach in financial trading, enabling agents to learn optimal strategies through direct market interaction.
By Lin Li, Li Rong Wang, Hsuan Fu, Xiuyi Fan
arXiv:2606. 04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets.
By Damian Lebied\'z, Robert \'Slepaczuk
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
The paper proposes an adversarial reinforcement‑learning framework for market making that incorporates Hawkes‑process driven order arrivals and trade‑induced price impact, addressing limitations of prior Poisson‑based models. An LSTM module captures temporal dependencies in recent observations to handle increased non‑stationarity, and the authors analyze equilibrium properties and introduce a robustness evaluation protocol focused on the left tail of returns. Experiments across diverse market regimes demonstrate that the method improves left‑tail performance, especially under strong Hawkes excitation and moderate price impact, without relying on a terminal inventory bias.
By Hao Yang, Zhenguo Xu