arXiv:2505.13299v2 Announce Type: replace-cross
Abstract: This paper considers the estimation of quantiles via a smoothed version of the stochastic gradient descent (SGD) algorithm. By smoothing the...
By Likai Chen, Georg Keilbar, Wei Biao Wu
arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv:2608. 27313v2 Announce Type: replace-cross Abstract: Quantile temporal-difference learning (QTD) is an effective method for learning return distributions through quantile approximation, yet its finite-time behavior remains poorly understood.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning. The proof separates two stability mechanisms.
arXiv:2603. 06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards.
By Alireza Mousavi-Hosseini, Murat A. Erdogdu
arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.
By Yang Peng, Liangyu Zhang
arXiv:2608. 27313v1 Announce Type: cross Abstract: We establish a global finite-sample guarantee for synchronous quantile temporal-difference learning (QTD) in tabular distributional reinforcement learning.
By Zijie Cheng, Xiang Li, Yang Peng, Zhihua Zhang
arXiv:2501.06926v5 Announce Type: replace
Abstract: Double reinforcement learning (DRL) provides efficient off-policy inference for policy values in nonparametric Markov decision processes (MDPs), bu...
By Lars van der Laan, David Hubbard, Allen Tran, Nathan Kallus, Aur\'{e}lien Bibaut
arXiv:2608. 08204v1 Announce Type: cross Abstract: This work proposes deep nonparametric Instrumental variable quantile regression (IVQR), a two-stage estimator that combines conditional diffusion modeling with a kernel-smoothed conditional moment formulation.
By Xingdong Feng, Xinhong Jiang, Yuling Jiao, Lican Kang, Junwei Liu
arXiv:2602. 09300v2 Announce Type: replace Abstract: We consider the policy evaluation and control in a finite horizon reinforcement learning (RL) setting under an expectile-based objective.
By Shrey Rakeshkumar Patel, Sumedh Gupte, Soumen Pachal, Prashanth L. A., Sanjay P. Bhat
arXiv:2607. 04627v1 Announce Type: new Abstract: Persona-Trained Monte Carlo (PTMC) estimates distributions of market-outcome functionals by repeatedly simulating limit-order-book interaction among $K$ neural policy bots whose behavioral personas are drawn from a learned heterogeneity distribution $\mathcal{P}$.
By Salavat Ishbulatov
Minimax risk and regret are expectation-based criteria and do not capture rare but consequential failures. To address this concern, we develop a $δ$-explicit minimax-quantile theory for interactive statistical decision making (ISDM).