arXiv:2404. 13879v5 Announce Type: replace Abstract: Uncertainties in transition dynamics pose a critical challenge in reinforcement learning (RL), often resulting in performance degradation of trained policies when deployed on hardware.
By Xulin Chen, Ruipeng Liu, Zhenyu Gan, Garrett E. Katz
In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been propo...
arXiv:2608.24146v1 Announce Type: new
Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate...
By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
By Yang Xu, Swetha Ganesh, Vaneet Aggarwal
The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
arXiv:2606. 00680v1 Announce Type: new Abstract: Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets.
By Hongqiang Lin, Pengfei Wang, Nenggan Zheng
arXiv:2607. 26509v1 Announce Type: new Abstract: Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement.
By Gong Gao, Xiao Lai, Ziqi Xie, Guojie Chen, Xianhui Liu, Weidong Zhao
The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.
By Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote, Javier Segovia-Aguas
arXiv:2510. 01721v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data.
By Saptarshi Mandal, Yashaswini Murthy, R. Srikant
arXiv:2608. 10634v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL), which learns environment dynamics to generate synthetic experience, is a promising approach to sample-efficient decision making.
By Zefeng Liang, Jie Qiao, Ruichu Cai, Weilin Chen, Zhifeng Hao
The paper introduces a model-based bootstrap framework for uncertainty quantification in offline policy evaluation (OPE) within finite-horizon, time-inhomogeneous Markov decision processes. Unlike traditional bootstrap methods that resample entire episodes, this approach regenerates trajectories from an estimated MDP, enabling use of diverse offline data formats such as complete trajectories, transition-level observations, and trajectory fragments. The authors prove bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation, and demonstrate through simulations that the method yields tighter confidence intervals and more accurate variance estimates compared to existing techniques.
By Weiwei Wang, Yuqiang Li, Xianyi Wu, Bingyi Jing
The paper introduces Bidirectional Behavior Prior Distillation (B2PD), a method that uses action‑value priors to train a conditional variational autoencoder for generating high‑value behavior support. These expert behavior priors are then distilled into the online reinforcement learning agent, reducing inefficient exploration and stabilizing policy updates. Experiments on state‑ and pixel‑based tasks show that B2PD improves sample efficiency while maintaining stable learning dynamics.
By Gong Gao, Xiao Lai, Jiaji Shen, Ning Jia, Xianhui Liu, Weidong Zhao