arXiv:2504. 10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies.
By Lukas-Benedikt Fiechtner, Jose Blanchet
arXiv:2605. 00155v3 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility.
By Yikai Wang, Shang Liu, Jose Blanchet
arXiv:2606. 19117v1 Announce Type: cross Abstract: Offline policy learning has received growing attention in causal inference.
By Yiyan Huang, Cheuk Hang Leung, Qi Wu, Zhiheng Zhang
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2502. 17602v2 Announce Type: replace-cross Abstract: We study a class of stochastic nonsmooth optimization problems in which an outer variable minimizes the expectation of a pointwise maximum.
By Wei Liu, Muhammad Khan, Gabriel Mancino-Ball, Yangyang Xu
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).
By Yang Xu, Swetha Ganesh, Vaneet Aggarwal