arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
arXiv:2605. 16103v2 Announce Type: replace Abstract: Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, positive errors can be selected and propagated, causing learned values to exceed the true optimal values.
By Donghwan Lee
arXiv:2505.13299v2 Announce Type: replace-cross
Abstract: This paper considers the estimation of quantiles via a smoothed version of the stochastic gradient descent (SGD) algorithm. By smoothing the...
By Likai Chen, Georg Keilbar, Wei Biao Wu
arXiv:2607. 08444v1 Announce Type: cross Abstract: In this paper, we study quantile-based distributional reinforcement learning from the perspective of statistical efficiency.
By Zijie Cheng, Yang Peng, Zhihua Zhang
arXiv:2608. 12973v1 Announce Type: cross Abstract: In this paper, we study how to perform statistical inference for quantile temporal difference learning (QTD) in distributional reinforcement learning.
By Zijie Cheng, Yang Peng, Zhihua Zhang