arXiv Machine Learning

Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples

arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).

Hugging Face Trending Papers
Sep 24

A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning

The paper introduces a contraction framework for stochastic operators that incorporates bootstrapping, where a variable is updated using a frozen copy as a target that is refreshed every $K$ steps. By modeling the sampled update as a stochastic operator, the authors derive a finite‑time bound for i.i.d. samples that applies to any target‑update period and does not require gradient structure or uniformly bounded sampling error. The framework shows that the iterates converge geometrically in root mean square to a ball around the fixed point, with the error floor scaling with the step size, and it generalizes existing deterministic and stochastic‑gradient bounds.