Model-based Bootstrap for Offline Policy Evaluation in Tabular Reinforcement Learning
Read the original on arXiv Machine Learning →The paper introduces a model-based bootstrap framework for uncertainty quantification in offline policy evaluation (OPE) within finite-horizon, time-inhomogeneous Markov decision processes. Unlike traditional bootstrap methods that resample entire episodes, this approach regenerates trajectories from an estimated MDP, enabling use of diverse offline data formats such as complete trajectories, transition-level observations, and trajectory fragments. The authors prove bootstrap distributional consistency, asymptotically valid confidence intervals, and consistent variance estimation, and demonstrate through simulations that the method yields tighter confidence intervals and more accurate variance estimates compared to existing techniques.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.