arXiv:2608. 12973v1 Announce Type: cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning.
By Zijie Cheng, Yang Peng, Zhihua Zhang
arXiv:2605. 12410v2 Announce Type: replace-cross Abstract: We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown.
By Ziwei Su, Imon Banerjee, Diego Klabjan
arXiv:2607. 08444v1 Announce Type: cross Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency.
By Zijie Cheng, Yang Peng, Zhihua Zhang
arXiv:2609.14922v1 Announce Type: cross
Abstract: For constant-stepsize stochastic approximation (SA), the iterates converge in distribution to a stationary law that depends on the stepsize $\alpha.$...
By Yixuan Zhang, Qiaomin Xie
arXiv:2505.13299v2 Announce Type: replace-cross
Abstract: This paper considers the estimation of quantiles via a smoothed version of the stochastic gradient descent (SGD) algorithm. By smoothing the...
By Likai Chen, Georg Keilbar, Wei Biao Wu
arXiv:2602.13960v2 Announce Type: replace
Abstract: Constant-stepsize stochastic approximation (SA) is widely used in learning for computational efficiency, yet the distribution of the iterates is ty...
By Zedong Wang, Yuyang Wang, Ijay Narang, Felix Wang, Yuzhou Wang, Siva Theja Maguluri
The paper investigates the finite‑iteration behavior of exact asynchronous recursions used in categorical distributional temporal‑difference (TD) learning. It analyzes both scalar categorical TD in the Cramér geometry and multivariate signed‑categorical TD in the maximum mean discrepancy geometry, showing that these methods can be viewed as single‑state stochastic‑approximation recursions that contract in a block‑supremum norm. The authors develop a restricted‑domain theory, derive discounted bounds under i.i.d. and Markovian sampling, and extend the analysis to undiscounted fixed‑horizon policy evaluation with horizon‑stacked categorical methods under episodic sampling, thereby providing a unified non‑asymptotic analysis across various settings.
By Ege C. Kaya, Abolfazl Hashemi
Suppose we observe the first $n$ points of a sequence of random variables having length $n+1$, and wish to estimate a functional of the unobserved final point and the empirical measure of the $n$ obse...
The paper introduces a leave‑a‑window‑out estimator for next‑token functionals, such as the surprise probability and test error, in sequences of random variables. By deleting a window of length τ after each index, the estimator generalizes leave‑one‑out and achieves parametric error decay for stationary β‑mixing processes that admit a Marton coupling. The authors provide both upper bounds and a minimax lower bound for the surprise probability, and demonstrate through simulations that their method outperforms traditional baselines on Markov, moving‑average, and autoregressive processes.
By Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
arXiv:2605. 26000v2 Announce Type: replace-cross Abstract: Stochastic gradient descent (SGD) is foundational to large-scale statistical learning and stochastic optimization.
By Jose Blanchet, Peter Glynn, Wenhao Yang
arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang