← Back to all news
arXiv Statistics ML October 1, 2026 By Yang Peng

Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data

Read the original on arXiv Statistics ML →

The Flow has not summarised this story yet — read it at arXiv Statistics ML.

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

Hugging Face Trending Papers
4d ago

Sharp Statistical Rates for Asynchronous TD Learning with Markovian Data

We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $γ$, write $H=(1-γ)^{-1}$, and let $μ_{\...

More like this →
arXiv Machine Learning
Sep 22

Optimal Sample Complexity of Stable Discounted Markov Decision Processes

arXiv:2302.07477v4 Announce Type: replace Abstract: We study the optimal sample complexity of tabular reinforcement learning for infinite-horizon discounted Markov decision processes. The unrestricte...

By Shengbo Wang, Jose Blanchet, Peter Glynn
reinforcement-learning
More like this →
arXiv Machine Learning
Jun 25

Minimax PAC Bounds for Learning in Exogenous Contextual MDPs

arXiv:2606. 25170v1 Announce Type: cross Abstract: We study PAC learning in tabular discounted Markov decision processes with exogenous i.

By Corentin Pla, Hugo Richard, Marc Abeille, Vianney Perchet
agentsreinforcement-learning
More like this →
Hugging Face Trending Papers
Aug 20

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.

reinforcement-learning
More like this →
arXiv Machine Learning
Jun 25

A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging

arXiv:2606. 24981v1 Announce Type: new Abstract: We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory.

By Wei-Cheng Lee, Francesco Orabona
More like this →
arXiv Machine Learning
Jul 23

Asymptotically Optimal Regret for Reinforcement Learning without Horizon Dependence

arXiv:2607. 19854v1 Announce Type: new Abstract: We study horizon-free regret minimization for finite-horizon time-homogeneous tabular Markov decision processes with $S$ states, $A$ actions, horizon $H$, and per-trajectory total reward bounded by $1$.

By Runlong Zhou, Zihan Zhang, Maryam Fazel, Simon S. Du
reinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea