arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).
By Ziad Kobeissi (L2S), \'Elo\"ise Berthier (U2IS)
arXiv:2506. 01052v3 Announce Type: replace Abstract: We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning.
By Wei-Cheng Lee, Francesco Orabona
arXiv:2609.14922v1 Announce Type: cross
Abstract: For constant-stepsize stochastic approximation (SA), the iterates converge in distribution to a stationary law that depends on the stepsize $\alpha.$...
By Yixuan Zhang, Qiaomin Xie
arXiv:2609.38880v1 Announce Type: new
Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discoun...
By Yang Peng
arXiv:2609.36859v1 Announce Type: new
Abstract: We identify and study a structural mechanism for Markovian nonconvex ADMM in reinforcement learning. Using finite discounted MDPs as a canonical provin...
By Zhaojun Peng
arXiv:2606. 24981v1 Announce Type: new Abstract: We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory.
By Wei-Cheng Lee, Francesco Orabona