arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.
By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
The paper introduces a model-based bootstrap framework for uncertainty quantification in offline policy evaluation (OPE) within finite-horizon, time-inhomogeneous Markov decision processes. Unlike traditional bootstrap methods that resample entire episodes, this approach regenerates trajectories from an estimated MDP, enabling use of diverse offline data formats such as complete trajectories, transition-level observations, and trajectory fragments. The authors prove bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation, and demonstrate through simulations that the method yields tighter confidence intervals and more accurate variance estimates compared to existing techniques.
By Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing
arXiv:2607. 23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning.
By Weikai Wang, Erick Delage
In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been propo...
arXiv:2608.24146v1 Announce Type: new
Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate...
By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang
The paper investigates continuous‑time stochastic control problems with unknown drift and running reward functions, using an exploratory reinforcement learning framework that incorporates relaxed controls and entropy regularization. It develops policy‑iteration algorithms based on probabilistic representations of the optimal value function and its gradient, proving convergence and demonstrating performance through numerical examples. The study also extends to a special case with control‑dependent diffusion, requiring a Hessian representation.
By Jin Ma, Gaozhan Wang, Jianfeng Zhang, Xunyu Zhou
arXiv:2408. 02295v4 Announce Type: replace Abstract: Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian.
By Seyeon Kim, Joonhun Lee, Namhoon Cho, Sungjun Han, Wooseop Hwang
arXiv:2606. 18186v1 Announce Type: cross Abstract: Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (when deployed on physical systems).
By Lekan Molu
arXiv:2606. 20206v1 Announce Type: cross Abstract: In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values.
By Ziheng Wei, Annie Qu, Rui Miao
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations rel...
arXiv:2607. 14522v1 Announce Type: new Abstract: We formulate reinforcement learning (RL) in continuous time with discrete state spaces and possibly arbitrary action spaces via a stochastic control approach, where the state dynamics are modeled as a controlled continuous-time Markov chain (CTMC).
By Zikun Zhang, Jiayuan Sheng, David D. Yao, Wenpin Tang
The paper investigates best‑policy identification in finite‑horizon, risk‑sensitive reinforcement learning using the entropic risk measure. It identifies a gap between known lower bounds ≥ η(e^{|eta|H}) and upper bounds ≤ O(e^{2|eta|H}) for sample complexity, attributing the excess factor to loose concentration bounds for exponential utilities. By employing a forward‑model algorithm with KL‑based exploration bonuses and a novel stopping rule, the authors achieve a sample complexity that matches the lower bound, closing the previously open exponential gap.
By Amer Essakine, Claire Vernade