We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.
By Yang Peng, Liangyu Zhang
arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
arXiv:2608. 27313v2 Announce Type: replace-cross Abstract: Quantile temporal-difference learning (QTD) is an effective method for learning return distributions through quantile approximation, yet its finite-time behavior remains poorly understood.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
arXiv:2505.13299v2 Announce Type: replace-cross
Abstract: This paper considers the estimation of quantiles via a smoothed version of the stochastic gradient descent (SGD) algorithm. By smoothing the...
By Likai Chen, Georg Keilbar, Wei Biao Wu
arXiv:2607. 08444v1 Announce Type: cross Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency.
By Zijie Cheng, Yang Peng, Zhihua Zhang