arXiv:2609.39837v1 Announce Type: new
Abstract: Policy mirror descent (PMD) enjoys fast convergence in regularized Markov decision processes (MDPs), but existing guarantees often rely on exact or inc...
By Qipei Chen, Wenye Li, Yule Sun, Ke Wei
We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discount factor $γ$, write $H=(1-γ)^{-1}$, and let $μ_{\...
arXiv:2609.38880v1 Announce Type: new
Abstract: We study the last iterate of standard tabular temporal-difference (TD) learning from a single trajectory of a finite Markov reward process. For discoun...
By Yang Peng
arXiv:2503. 14549v3 Announce Type: replace-cross Abstract: How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a prescribed Gibbs distribution?
By Michael Chertkov, Sungsoo Ahn, Hamidreza Behjoo
Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning. A recent work \citep{thoppe2026reinforcement} addressed this difficulty by introducing a Bellman-compatible surrogate and two model-free fixed-point algorithms for optimizing it over stationary policies.
arXiv:2511.08097v2 Announce Type: replace-cross
Abstract: We consider a general infinite horizon Heterogeneous Restless multi-armed Bandit (RMAB). Heterogeneity is a fundamental problem for many real...
By Dheeraj Narasimha, Nicolas Gast
Limiting‑Kernel Q(λ) (LKQL) is an off‑policy value estimator that blends n‑step truncation with a long‑horizon approximation based on the limiting kernel. It maintains the computational efficiency of n‑step methods while improving policy evaluation accuracy, especially for long‑horizon tasks. The authors prove faster convergence of LKQL’s operator under aperiodicity and near‑on‑policy conditions, and demonstrate empirical gains on MuJoCo continuous‑control benchmarks.
By Tolga Ok, Arman Sharifi Kolarijani, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi Kolarijani
arXiv:2606. 15978v1 Announce Type: new Abstract: Tsitsiklis proved convergence of Monte Carlo optimistic policy iteration under a uniform update structure and identified nonuniform update frequencies as a delicate obstruction.
By Yuanlong Chen
arXiv:2606. 15247v1 Announce Type: cross Abstract: The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting.
By Octave Oliviers, Glenn Vinnicombe
The paper investigates the finite‑iteration behavior of exact asynchronous recursions used in categorical distributional temporal‑difference (TD) learning. It analyzes both scalar categorical TD in the Cramér geometry and multivariate signed‑categorical TD in the maximum mean discrepancy geometry, showing that these methods can be viewed as single‑state stochastic‑approximation recursions that contract in a block‑supremum norm. The authors develop a restricted‑domain theory, derive discounted bounds under i.i.d. and Markovian sampling, and extend the analysis to undiscounted fixed‑horizon policy evaluation with horizon‑stacked categorical methods under episodic sampling, thereby providing a unified non‑asymptotic analysis across various settings.
By Ege C. Kaya, Abolfazl Hashemi
arXiv:2608. 01917v1 Announce Type: new Abstract: Discounted exponential utility provides a principled criterion for risk-sensitive sequential decision-making, but its nonlinear structure complicates reinforcement learning.
By Ankur Naskar, Vivek T A, Aditya Kumar, Gugan Thoppe, Prashanth L. A
arXiv:2607. 02137v1 Announce Type: cross Abstract: We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid.
By Yilie Huang, Wenpin Tang, Xun Yu Zhou