arXiv Machine Learning

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

arXiv:2607. 23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning.

arXiv Statistics ML
Sep 25

Shrinking-Tube Concentration for Adaptive Markovian Stochastic Approximation

The paper establishes a shrinking‑tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain, guaranteeing that after a chosen time every iterate stays within a tolerance that tightens over time. The bound’s probability of any exit after that time decays polynomially, and a matching lower bound shows this exponent is optimal under finite second moments. Extensions to recursions with martingale‑difference noise and predictable bias reveal how noise scale and bias affect exit‑probability decay and tube shrinkage, with applications to inventory learning and numerical gradient accuracy.

By Jin Li, Ye Luo, Xiaowei Zhang
arXiv Machine Learning
1d ago

Towards Optimal Policy Improvement

The paper introduces a framework for optimal policy improvement in reinforcement learning, defining it as the best single update under given constraints. It shows that restricting improvement to a subset of states is equivalent to solving an induced Markov Decision Process, linking planning with explicit or implicit models to optimal policy improvement. The authors develop a novel operator for greedification under approximate evaluation, demonstrating empirical gains across several RL algorithms and settings.

By Yaniv Oren, Viliam Vadocz, Wiktor Zabka, Thomas Evers, Jan Robine, Wendelin B\"ohmer, Matthijs T. J. Spaan, Martha White, Hendrik Baier, Fenghui Yu
arXiv Machine Learning
Jun 24

A Robust Model-Based Approach for Continuous-Time Policy Evaluation with Unknown L\'evy Process Dynamics

arXiv:2504. 01482v3 Announce Type: replace-cross Abstract: This paper develops a model-based framework for continuous-time policy evaluation (CTPE) in reinforcement learning, incorporating both Brownian and L\'evy noise to model stochastic dynamics influenced by rare and extreme events.

By Qihao Ye, Xiaochuan Tian, Yuhua Zhu