arXiv:2609.14922v1 Announce Type: cross
Abstract: For constant-stepsize stochastic approximation (SA), the iterates converge in distribution to a stationary law that depends on the stepsize $\alpha.$...
By Yixuan Zhang, Qiaomin Xie
arXiv:2506. 01052v3 Announce Type: replace Abstract: We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning.
By Wei-Cheng Lee, Francesco Orabona
arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.
By Asha Barua, Sajad Khodadadian
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $γ$, write $H=(1-γ)^{-1}$, and let $μ_{\...
arXiv:2602.13960v2 Announce Type: replace
Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is ty...
By Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
arXiv:2608. 19643v1 Announce Type: new Abstract: Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses.
By Yi-Shan Wu
The paper studies least squares parameter estimation for discrete‑time, unstable, closed‑loop nonlinear stochastic systems with linearly parametrised uncertainty and additive i.i.d. process noise. By perturbing the control policy with exploratory input and assuming a sub‑exponential input‑to‑state growth property, the authors derive non‑asymptotic bounds on the estimation error whenever the state trajectory remains in an informative region of the state space. When the entire state space is informative, the bounds hold with high probability for all time steps, and the authors illustrate the applicability of their results with examples that extend beyond existing work.
By Seth Siriya, Jingge Zhu, Dragan Ne\v{s}i\'c, Ye Pu
Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems.
arXiv:2606. 05967v1 Announce Type: cross Abstract: In this paper, we study the finite-time behavior of the TD(0) temporal-difference method with linear function approximation (LFA).
By Ziad Kobeissi (L2S), \'Elo\"ise Berthier (U2IS)