A Contraction Framework for Stochastic Operators with Bootstrapping: Application to TD Learning
arXiv:2609. 29961v1 Announce Type: new Abstract: Many iterative algorithms rely on bootstrapping.
The paper introduces a contraction framework for stochastic operators that incorporates bootstrapping, where a variable is updated using a frozen copy as a target that is refreshed every $K$ steps. By modeling the sampled update as a stochastic operator, the authors derive a finite‑time bound for i.i.d. samples that applies to any target‑update period and does not require gradient structure or uniformly bounded sampling error. The framework shows that the iterates converge geometrically in root mean square to a ball around the fixed point, with the error floor scaling with the step size, and it generalizes existing deterministic and stochastic‑gradient bounds.
arXiv:2609. 29961v1 Announce Type: new Abstract: Many iterative algorithms rely on bootstrapping.
arXiv:2505.01361v3 Announce Type: replace Abstract: Temporal difference (TD) learning is a foundational algorithm in reinforcement learning (RL). For nearly forty years, TD learning has served as a w...
arXiv:2506. 01052v3 Announce Type: replace Abstract: We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning.
arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).
The paper introduces Batched SGD, a variant that groups online samples into epochs and performs a single update per epoch using a low‑variance gradient estimate. This batching approach allows a straightforward high‑probability analysis without restrictive assumptions or auxiliary sequences, yielding near‑optimal rates for both strongly convex and non‑convex objectives under standard smoothness and sub‑Gaussian noise conditions. The authors also extend the method to federated learning, providing the first high‑probability guarantees with logarithmic communication complexity, linear speedup in the number of agents, and robustness to data heterogeneity.
arXiv:2607. 17595v1 Announce Type: new Abstract: We establish mean-square and concentration bounds for stochastic approximation (SA) with arbitrary norm contractive mappings, under a multiplicative noise model where the noise may scale affinely with the norm of the iterates, and the iterates are potentially unbounded.
arXiv:2609.39837v1 Announce Type: new Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA). We consider on-policy independent and identically distributed (i.
arXiv:2608. 25551v1 Announce Type: new Abstract: Stochastic gradient descent (SGD) is typically analyzed at a deterministic horizon chosen before the algorithm is run, even though practical stopping decisions are made adaptively by inspecting the evolving trajectory.
arXiv:2504.18184v5 Announce Type: replace Abstract: We consider a class of statistical inverse problems involving the estimation of a regression operator from a Polish space to a separable Hilbert sp...
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2604. 06039v2 Announce Type: replace-cross Abstract: Value iteration-type methods have been extensively studied for computing a nearly optimal value function in reinforcement learning (RL).