arXiv:2606. 04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets.
By Damian Lebied\'z, Robert \'Slepaczuk
arXiv:2607. 28127v1 Announce Type: cross Abstract: Recent advances in Generative AI have substantially improved financial sentiment analysis through post-trained financial large language models (LLMs).
By Giorgos Iacovides, Wuyang Zhou, Danilo Mandic
arXiv:2609.13825v1 Announce Type: new
Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) o...
By Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita
arXiv:2606. 27032v1 Announce Type: cross Abstract: Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints.
By Jesper Klicks, Sander Vr\v{z}ina, Vincent Fran\c{c}ois-Lavet
arXiv:2606. 08450v1 Announce Type: new Abstract: Financial portfolio trading is naturally formulated as a reinforcement learning problem, where an agent sequentially rebalances assets under changing market conditions to balance return, risk, and transaction costs.
By Yanyan Wu, Boyi Zhang, Yanlin Liu, Xinyu Fang, Jining Luan, Meiqi Zhang, Jiacheng Liu, Hao Zeng, Dexu Yu, Chang Liu, Hanwen Du, Yongxin Ni, Youhua Li
arXiv:2606. 08379v1 Announce Type: new Abstract: This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smoothing), a deterministic actor-critic architecture that combines twin exponential-moving-average critic targets with pessimistic min backup, TD3-style target policy smoothing noise, delayed actor updates, and conservative Q regularisation to curb overestimation.
By Ilia Zaznov, Atta Badii, Julian Kunkel, Alfonso Dufour
arXiv:2606. 30997v1 Announce Type: new Abstract: We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared by all prior financial RL work: 1) ticker lock-in, 2) monolithic objectives , and 3) static user models.
By Ramin Pishehvar
The paper introduces an agentic forecasting environment built on 2,100+ resolved Polymarket questions, where a language model (Qwen3.5-35B-A3B) learns to gather evidence during rollout via web search, page reading, and financial time series, all filtered to avoid post‑cutoff leaks. Training with single‑epoch GRPO and a Brier‑score reward improves calibration by 30‑40% and reduces search attempts, while the trained policy outperforms four frontier models in evidence‑based forecasting, achieving lower soft‑Brier scores at roughly 5% of the inference cost. The authors release the environment, dataset, and per‑rollout records as a reusable harness for temporal forecasting agents.
By Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha
arXiv:2608. 04305v1 Announce Type: new Abstract: Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and inefficient sample reuse.
By Yifan Wu, Junjie Lei, Wenjie Huang
arXiv:2606. 00060v1 Announce Type: cross Abstract: This paper investigates whether machine learning forecasts of hourly BTC-USDT returns can be converted into economically meaningful trading performance after transaction costs.
By Andrei Bysik, Robert \'Slepaczuk
arXiv:2608. 15841v1 Announce Type: new Abstract: Reinforcement learning has gained increasing attention as a data-driven approach for stock trading.
By Arishi Orra, Himanshu Choudhary, Manoj Thakur
PPO‑HRAP introduces a hybrid regime‑aware policy that blends Proximal Policy Optimization with a volatility‑conditioned regime prior to balance upside participation and drawdown control in trading. The agent uses market and portfolio features, rewards that combine log return, VIX‑conditioned drawdown penalty, exposure deviation, and turnover cost, and outputs a blended action between the PPO actor and the regime‑derived target exposure. In backtests on SPY (2020‑2022) it achieved a 27.62% total return, 8.48% annualized return, and reduced maximum drawdown from 34.10% to 18.47%, while maintaining stable performance across multiple seeds and ranking first on total return and Sharpe ratio in single‑run cross‑asset tests on QQQ and DIA.
By Duong Hien Chi Kien, Thanh Trung Huynh