arXiv:2607. 16028v1 Announce Type: new Abstract: This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data.
By Andrei Neagu, Eeham Khan, Leila Kosseim
PPO‑HRAP introduces a hybrid regime‑aware policy that blends Proximal Policy Optimization with a volatility‑conditioned regime prior to balance upside participation and drawdown control in trading. The agent uses market and portfolio features, rewards that combine log return, VIX‑conditioned drawdown penalty, exposure deviation, and turnover cost, and outputs a blended action between the PPO actor and the regime‑derived target exposure. In backtests on SPY (2020‑2022) it achieved a 27.62% total return, 8.48% annualized return, and reduced maximum drawdown from 34.10% to 18.47%, while maintaining stable performance across multiple seeds and ranking first on total return and Sharpe ratio in single‑run cross‑asset tests on QQQ and DIA.
By Duong Hien Chi Kien, Thanh Trung Huynh
arXiv:2606. 04574v1 Announce Type: new Abstract: This study aims to determine whether the application of Deep Reinforcement Learning (DRL) as a specialized execution overlay can enhance pair trading in highly volatile cryptocurrency markets.
By Damian Lebied\'z, Robert \'Slepaczuk
arXiv:2606. 27032v1 Announce Type: cross Abstract: Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constraints.
By Jesper Klicks, Sander Vr\v{z}ina, Vincent Fran\c{c}ois-Lavet
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
By Dongsheng Hou, Yanqiao Chen, Yuhan Rui
arXiv:2609.13825v1 Announce Type: new
Abstract: Reinforcement learning trading systems published in the academic literature overwhelmingly rely on price-aggregate state representations (OHLCV bars) o...
By Asser Moustafa, Rares-Mihail Neagu, Jugal Kalita
arXiv:2609.36802v1 Announce Type: new
Abstract: A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning...
By Xuanyi Zhou, Qiuyang Mang, Huanzhi Mao, Dacheng Li, Wenhao Chai, Mayank Mishra, Yichuan Wang, Karthik Narasimhan, Alvin Cheung, Joseph E. Gonzalez
arXiv:2609.14327v1 Announce Type: new
Abstract: Variance penalization is a principled approach to risk-sensitive reinforcement learning (RL) that explicitly trades expected return for policy stabilit...
By Saunak Kumar Panda, Tong Li, Yisha Xiang, Ruiqi Liu
arXiv:2606. 30316v1 Announce Type: new Abstract: This paper studies Reinforcement Learning as an online controller for curtailment-aware workload shifting in wind-turbine-integrated high-performance computing (HPC) data centers.
By Jan Stenner, Alexander Kilian, Sebastian Peitz, Hermann de Meer
arXiv:2608. 02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization.
By Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
arXiv:2606. 23977v1 Announce Type: new Abstract: Efficient sorter diversion control of automated material handling systems (MHS) is critical for optimizing operational efficiency in large-scale warehouse environments.
By Tina Dongxu Li, Mouhacine Benosman, Ken Meszaros, Trevor Dardik
arXiv:2603. 21180v4 Announce Type: replace Abstract: Sequential experimental design under expensive, gradient-free objectives is a central challenge in computational statistics: evaluation budgets are tightly constrained and information must be extracted efficiently from each observation.
By Foo Hui-Mean, Yuan-chin I Chang