arXiv Machine Learning By Weikai Wang, Erick Delage

Online Policy Evaluation for MDPs with Dynamic UBSR Measures

Read the original on arXiv Machine Learning →

arXiv:2607. 23030v1 Announce Type: new Abstract: Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
Sep 25

Shrinking-Tube Concentration for Adaptive Markovian Stochastic Approximation

The paper establishes a shrinking‑tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain, guaranteeing that after a chosen time every iterate stays within a tolerance that tightens over time. The bound’s probability of any exit after that time decays polynomially, and a matching lower bound shows this exponent is optimal under finite second moments. Extensions to recursions with martingale‑difference noise and predictable bias reveal how noise scale and bias affect exit‑probability decay and tube shrinkage, with applications to inventory learning and numerical gradient accuracy.

By Jin Li, Ye Luo, Xiaowei Zhang