The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.
By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes
The paper investigates model‑free robust Q‑learning with χ² uncertainty sets and linear function approximation, using data from a single trajectory of an unknown nominal MDP. It introduces a variational reformulation of the robust Bellman target and a blockwise frozen‑target scheme to overcome estimation and non‑contractivity challenges, and proves a finite‑time error bound for every discount factor γ in (0,1). A neural‑network experiment demonstrates the practical use of the variational target in a continuous‑state nonlinear‑control task.
By Saptarshi Mandal, Yashaswini Murthy, R. Srikant
arXiv:2510. 01721v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DRRL) seeks policies that perform well when the deployment transition model differs from the nominal model generating the data.
By Saptarshi Mandal, Yashaswini Murthy, R. Srikant
arXiv:2603. 09344v3 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often degrades under distribution shift.
By Hongqiang Lin, Zhenghui Fu, Weihao Tang, Pengfei Wang, Yiding Sun, Qixian Huang, Dongxu Zhang
arXiv:2506. 01052v3 Announce Type: replace Abstract: We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning.
By Wei-Cheng Lee, Francesco Orabona
The paper introduces a vector Bellman theory for multichain robust average‑reward Markov decision processes, addressing the state‑dependent optimal long‑run rewards that arise under uncertainty. It develops a gain‑first, bias‑second optimization principle for finite models with compact, post‑action $(s,a)$‑rectangular ambiguity, yielding a coupled vector gain‑bias system and stationary saddle strategies from all initial states. The authors also characterize solvability conditions, provide certificates for asymptotically affine trajectories of the robust Bellman operator, and design a robust approximately shifted Halpern planning algorithm that converges to the optimal gain vector and produces average‑optimal greedy controllers.
By Yue Wang, George Atia
Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.
By Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani
arXiv:2602. 03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events.
By Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon, Erick Delage
arXiv:2608.22636v1 Announce Type: cross
Abstract: Q-learning with linear function approximation can be unstable because an arbitrary approximation architecture need not preserve the Bellman contracti...
By Shengbo Wang
arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
In value-based reinforcement learning, improving the accuracy of policy evaluation has been shown to improve downstream policy optimization performance. The widely adopted family of approximations rel...
arXiv:2404. 03578v3 Announce Type: replace Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL).
By Miao Lu, Han Zhong, Tong Zhang, Jose Blanchet