arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv:2603. 06957v2 Announce Type: replace-cross Abstract: We study post-training linear autoregressive models with outcome and process rewards.
By Alireza Mousavi-Hosseini, Murat A. Erdogdu
arXiv:2608. 14408v1 Announce Type: cross Abstract: We study online statistical inference for functionals of the return distribution under a fixed policy.
By Yang Peng, Liangyu Zhang
arXiv:2608. 08204v1 Announce Type: cross Abstract: This work proposes deep nonparametric Instrumental variable quantile regression (IVQR), a two-stage estimator that combines conditional diffusion modeling with a kernel-smoothed conditional moment formulation.
By Xingdong Feng, Xinhong Jiang, Yuling Jiao, Lican Kang, Junwei Liu
arXiv:2607. 04627v1 Announce Type: new Abstract: Persona-Trained Monte Carlo (PTMC) estimates distributions of market-outcome functionals by repeatedly simulating limit-order-book interaction among $K$ neural policy bots whose behavioral personas are drawn from a learned heterogeneity distribution $\mathcal{P}$.
By Salavat Ishbulatov
Minimax risk and regret are expectation-based criteria and do not capture rare but consequential failures. To address this concern, we develop a $δ$-explicit minimax-quantile theory for interactive statistical decision making (ISDM).