The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.
By Deep Kumar Ganguly, Jan Kretinsky
arXiv:2609.00455v1 Announce Type: new
Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning cap...
By Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
By Zihan Zhu, Shayan Kiyani, George Pappas. Hamed Hassani
arXiv:2605. 23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.
By Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P\'erez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
arXiv:2607. 02206v1 Announce Type: cross Abstract: Predictions are increasingly used to guide high-stakes decisions, from treatment selection to policy making.
By Yurui Zheng, Ying Jin