arXiv AI

Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk

The paper introduces the Wasserstein entropic value-at-risk, a coherent risk measure that replaces the relative-entropy ball of the traditional entropic value-at-risk with an optimal-transport ball. This new measure captures reachable catastrophes that the original entropic measure ignores, and its variational dual mirrors the entropic formula with a transport price replacing inverse temperature. By driving the transport radius with belief entropy, the authors derive a closed‑form robust dynamic‑programming operator whose cautiousness decreases as belief sharpens, providing a certified safety sandwich and a sharp safety switch.

arXiv AI
Aug 19

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.

By Deep Kumar Ganguly, Jan Kretinsky
arXiv Machine Learning
Jun 30

Wasserstein Distributionally Robust Regret Optimization

arXiv:2504. 10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies.

By Lukas-Benedikt Fiechtner, Jose Blanchet
arXiv Machine Learning
Aug 12

Risk-Averse Wasserstein Distributionally Robust Online Learning

arXiv:2602. 20403v2 Announce Type: replace Abstract: We study distributionally robust online learning, where a risk-averse learner updates decisions sequentially to guard against worst-case distributions drawn from a Wasserstein ambiguity set centered at past observations.

By Guixian Chen, Salar Fattahi, Soroosh Shafiee
arXiv AI
Jul 7

Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control

arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.

By Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum
arXiv Machine Learning
Sep 17

Wasserstein Formulation of Reinforcement Learning. An Optimal Transport Perspective on Policy Optimization

The paper introduces a geometric framework for reinforcement learning that treats policies as mappings into the Wasserstein space of action probabilities. It establishes a Riemannian structure induced by stationary distributions, defines the tangent space of policies, and characterizes geodesics while addressing measurability concerns. The authors formulate a general RL optimization problem, construct a gradient flow via Otto's calculus, compute the gradient and Hessian of the energy, and demonstrate the approach with numerical examples for low‑dimensional problems and neural‑network‑parameterized policies for high‑dimensional settings.

By Mathias Dus (IRMA)
arXiv Machine Learning
Jun 11

Calibrating Decision Robustness via Inverse Conformal Risk Control

arXiv:2510. 07750v3 Announce Type: replace-cross Abstract: Robust optimization safeguards decisions against uncertainty by optimizing against worst-case scenarios, yet their effectiveness hinges on a prespecified robustness level that is often chosen ad hoc, leading to either insufficient protection or overly conservative and costly solutions.

By Wenbin Zhou, Shixiang Zhu