The paper introduces a penalized distributionally robust optimization framework that allows an adversary to choose any distribution while incurring a Wasserstein penalty for deviating from the empirical distribution. It shows that the adversary’s problem can be reformulated as optimizing transport maps that push empirical samples to adversarial ones, proving that optimal maps are cyclically monotone. The authors argue that standard per-sample adversarial training violates this property and propose two remedies—multi-start particle ascent and input-convex neural network parameterization—to enforce cyclical monotonicity, demonstrating improved robustness and generalization in experiments on regression, image classification, and control tasks.
By Alireza Abdollahpoorrostam, Ehsan Sharifian, Buse \c{S}en, Marco Cuturi, Daniel Kuhn
arXiv:2509.26364v3 Announce Type: replace
Abstract: The Schr\"odinger bridge problem is concerned with finding a stochastic dynamical system bridging two marginal distributions that minimises a certa...
By Kirill Tamogashev, Esmeralda S. Whitammer
arXiv:2605. 00155v3 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility.
By Yikai Wang, Shang Liu, Jose Blanchet
arXiv:2412. 20556v2 Announce Type: replace-cross Abstract: We study distributionally robust optimization (DRO) for robust inference when the worst-case distribution is continuous, leading to significant computational challenges due to the infinite-dimensional nature of the optimization problem.
By Linglingzhi Zhu, Yunqin Zhu, Yao Xie
arXiv:2404. 03578v3 Announce Type: replace Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL).
By Miao Lu, Han Zhong, Tong Zhang, Jose Blanchet
arXiv:2509. 09371v2 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) protects statistical learning against distributional shifts by optimizing the worst-case performance over a set of perturbed distributions.
By Zitao Wang, Nian Si, Molei Liu
arXiv:2608. 11544v1 Announce Type: cross Abstract: We propose CVaR-penalized Generative Particle Algorithm (CVaR-GPA), a robust, tail-agnostic algorithm for fine-tuning generative models to learn heavy-tailed distributions and capture extreme events, requiring no prior knowledge or estimation of the target's tail characteristics.
By Thejani Gamage, Hyemin Gu, Zhizhen Zhang, Ziyu Chen, Markos Katsoulakis, Luc Rey-Bellet
arXiv:2606. 15359v1 Announce Type: new Abstract: Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories.
By Paolo Giaretta, Zeyang Li, Navid Azizan
arXiv:2602. 21429v3 Announce Type: replace Abstract: Flow-based generative models, such as diffusion models and flow matching models, have achieved remarkable success in learning complex data distributions.
By Darshan Gadginmath, Ahmed Allibhoy, Fabio Pasqualetti
arXiv:2505. 07124v3 Announce Type: replace Abstract: We study inverse problems where an unknown potential is observed only through samples from the measure it induces by a convex variational principle.
By Francisco Andrade, Gabriel Peyr\'e, Clarice Poon
arXiv:2608.22746v1 Announce Type: new
Abstract: This paper studies the Sinkhorn distributionally robust hypothesis testing (SDRHT) problem, seeking a robust detector against least-favorable distribut...
By Fenglin Zhang, Teyan Liu, Jie Wang
arXiv:2606. 30230v1 Announce Type: cross Abstract: Learned reconstruction operators for inverse problems are typically trained under a fixed noise model, and generalize poorly when the distribution during testing differs from the one assumed during training.
By Floor van Maarschalkerwaart, Subhadip Mukherjee, Christoph Brune, Marcello Carioni