The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
arXiv:2607. 29593v1 Announce Type: new Abstract: This paper studies the policy gradient update for a multi-arm bandit problem in diffusion environment that is described by a stochastic differential equation (SDE) under the continuous-time reinforcement learning framework by Wang et al.
By Yanwei Jia, Du Ouyang
arXiv:2605. 16103v2 Announce Type: replace Abstract: Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, positive errors can be selected and propagated, causing learned values to exceed the true optimal values.
By Donghwan Lee
arXiv:2606. 16846v1 Announce Type: cross Abstract: We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions.
By Qian Qi
arXiv:2608. 02332v1 Announce Type: new Abstract: In offline reinforcement learning (RL), the distribution shift between behavioral data and the learned policy can lead to erroneous \emph{Q}-value estimation, thereby misguiding the direction of policy optimization.
By Botao Dong, Longyang Huang, Ning Pang, Hongtian Chen
arXiv:2510. 03494v2 Announce Type: replace Abstract: We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization.
By Volodymyr Tkachuk, Csaba Szepesv\'ari, Xiaoqi Tan