arXiv Machine Learning

CLaC@FinMMEval 2026 Task 3: Sentiment-Augmented Deep Reinforcement Learning for Active Trading -- An Alpha-Reward Approach

arXiv:2607. 16028v1 Announce Type: new Abstract: This paper presents our system for Task 3 of the CLEF 2026 FinMMEval Lab, which requires daily long, flat, or short trading decisions for Bitcoin (BTC) and Tesla (TSLA) using news and historical market data.

arXiv AI
Jun 9

GIFT: LLM-Guided State-Reward Interface for Financial Reinforcement Learning

arXiv:2606. 08450v1 Announce Type: new Abstract: Financial portfolio trading is naturally formulated as a reinforcement learning problem, where an agent sequentially rebalances assets under changing market conditions to balance return, risk, and transaction costs.

By Yanyan Wu, Boyi Zhang, Yanlin Liu, Xinyu Fang, Jining Luan, Meiqi Zhang, Jiacheng Liu, Hao Zeng, Dexu Yu, Chang Liu, Hanwen Du, Yongxin Ni, Youhua Li
arXiv AI
Jun 9

TT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution

arXiv:2606. 08379v1 Announce Type: new Abstract: This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smoothing), a deterministic actor-critic architecture that combines twin exponential-moving-average critic targets with pessimistic min backup, TD3-style target policy smoothing noise, delayed actor updates, and conservative Q regularisation to curb overestimation.

By Ilia Zaznov, Atta Badii, Julian Kunkel, Alfonso Dufour
arXiv Machine Learning
1d ago

Do Your Own Research: Learning to Forecast by Learning to Search

The paper introduces an agentic forecasting environment built on 2,100+ resolved Polymarket questions, where a language model (Qwen3.5-35B-A3B) learns to gather evidence during rollout via web search, page reading, and financial time series, all filtered to avoid post‑cutoff leaks. Training with single‑epoch GRPO and a Brier‑score reward improves calibration by 30‑40% and reduces search attempts, while the trained policy outperforms four frontier models in evidence‑based forecasting, achieving lower soft‑Brier scores at roughly 5% of the inference cost. The authors release the environment, dataset, and per‑rollout records as a reusable harness for temporal forecasting agents.

By Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha
arXiv Machine Learning
Aug 6

Adaptive Finite-Budget Training for CVaR Risk-Aware Q-Learning

arXiv:2608. 04305v1 Announce Type: new Abstract: Risk-aware Q-learning (RaQL) provides a model-free, two-timescale estimator for dynamic risk objectives, but its finite-budget behavior remains fragile: fixed inner-loop hyperparameters can produce unstable value estimates, persistent Bellman residuals, and inefficient sample reuse.

By Yifan Wu, Junjie Lei, Wenjie Huang
arXiv AI
2d ago

PPO-HRAP: Proximal Policy Optimization with a Hybrid Regime-Aware Policy for Risk-Controlled Trading

PPO‑HRAP introduces a hybrid regime‑aware policy that blends Proximal Policy Optimization with a volatility‑conditioned regime prior to balance upside participation and drawdown control in trading. The agent uses market and portfolio features, rewards that combine log return, VIX‑conditioned drawdown penalty, exposure deviation, and turnover cost, and outputs a blended action between the PPO actor and the regime‑derived target exposure. In backtests on SPY (2020‑2022) it achieved a 27.62% total return, 8.48% annualized return, and reduced maximum drawdown from 34.10% to 18.47%, while maintaining stable performance across multiple seeds and ranking first on total return and Sharpe ratio in single‑run cross‑asset tests on QQQ and DIA.

By Duong Hien Chi Kien, Thanh Trung Huynh