arXiv Machine Learning By Wei-Cheng Lee, Francesco Orabona

A Robust $\widetilde{\mathcal{O}}(1/\sqrt{T})$ Rate for Unprojected TD Learning with Linear Function Approximation

Read the original on arXiv Machine Learning →

arXiv:2506. 01052v3 Announce Type: replace Abstract: We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Sep 24

A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning

The paper introduces a contraction framework for stochastic operators that incorporates bootstrapping, where a variable is updated using a frozen copy as a target that is refreshed every $K$ steps. By modeling the sampled update as a stochastic operator, the authors derive a finite‑time bound for i.i.d. samples that applies to any target‑update period and does not require gradient structure or uniformly bounded sampling error. The framework shows that the iterates converge geometrically in root mean square to a ball around the fixed point, with the error floor scaling with the step size, and it generalizes existing deterministic and stochastic‑gradient bounds.