Decision Making Needs Uncertainty Quantification [Lecture Notes]
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
The paper introduces the Wasserstein entropic value-at-risk, a coherent risk measure that replaces the relative-entropy ball of the traditional entropic value-at-risk with an optimal-transport ball. This new measure captures reachable catastrophes that the original entropic measure ignores, and its variational dual mirrors the entropic formula with a transport price replacing inverse temperature. By driving the transport radius with belief entropy, the authors derive a closed‑form robust dynamic‑programming operator whose cautiousness decreases as belief sharpens, providing a certified safety sandwich and a sharp safety switch.
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
arXiv:2603. 02491v3 Announce Type: replace-cross Abstract: As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty?
arXiv:2609.38938v1 Announce Type: new Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies...
arXiv:2606. 04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep RL may demonstrate susceptibility to transition perturbations that result in unknown or unsafe behaviour.
arXiv:2608. 03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications.
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
arXiv:2609.24103v1 Announce Type: new Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first st...
arXiv:2601. 22993v4 Announce Type: replace Abstract: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
arXiv:2506. 07040v4 Announce Type: replace-cross Abstract: We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs).