arXiv:2609.38880v1 Announce Type: new
Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discoun...
By Yang Peng
arXiv:2302.07477v4 Announce Type: replace
Abstract: We study the optimal sample complexity of tabular reinforcement learning for infinite-horizon discounted Markov decision processes. The unrestricte...
By Shengbo Wang, Jose Blanchet, Peter Glynn
arXiv:2606. 25170v1 Announce Type: cross Abstract: We study PAC learning in tabular discounted Markov decision processes with exogenous i.
By Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet
arXiv:2606. 24981v1 Announce Type: new Abstract: We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory.
By Wei-Cheng Lee, Francesco Orabona
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.
By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
arXiv:2608. 19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored.
By Zhiqiang Tan
arXiv:2607. 22982v1 Announce Type: new Abstract: Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success.
By Asha Barua, Sajad Khodadadian
arXiv:2609.36486v1 Announce Type: new
Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M...
By Zijun Chen, Zihan Zhang
arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
arXiv:2608. 19643v1 Announce Type: new Abstract: Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses.
By Yi-Shan Wu
arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei