The paper introduces RATTL (Risk-Adversarial Total-Reward Learning), a framework that adjusts an agent’s caution based on epistemic uncertainty by using a Bayesian posterior over dynamics and a Wasserstein ambiguity set whose radius depends on that posterior. As evidence accumulates, the radius shrinks, smoothly transitioning the agent’s behavior from worst-case robustness to risk-neutral reward maximization. The authors prove a Safety Sandwich theorem showing RATTL’s value lies between the uninformed robust value and the full-knowledge optimum, and demonstrate the method on a binary-hazard example where the criterion reduces to Conditional Value-at-Risk.
By Deep Kumar Ganguly, Jan Kretinsky
arXiv:2609.00455v1 Announce Type: new
Abstract: Large language models (LLMs) are being used as policies for autonomous decision-making and planning in many domains. Despite their strong reasoning cap...
By Shubham Kumar, Harshit Kumar, Narendra Ahuja, Saurabh Jha
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
By Zihan Zhu, Shayan Kiyani, George Pappas. Hamed Hassani
arXiv:2605. 23146v3 Announce Type: replace-cross Abstract: Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.
By Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Syed Mahir Ahamed, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Marina P\'erez del Valle, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
arXiv:2607. 02206v1 Announce Type: cross Abstract: Predictions are increasingly used to guide high-stakes decisions, from treatment selection to policy making.
By Yurui Zheng, Ying Jin
arXiv:2607. 02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors.
By Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso
The paper introduces Retrospective World Modeling, a new paradigm for vision‑language‑model (VLM) agents that allows them to reason backward by estimating which action most likely caused a state transition. It proposes the Self‑Consistency Reward (SCR), an intrinsic signal that measures how well a policy action aligns with this retrospective explanation, providing dense transition‑level feedback. Experiments demonstrate that incorporating SCR improves policy robustness and generalization compared to purely prospective world‑modeling approaches.
By Yongjiang Liu, Jie Zhang, Haoyue Zhang, Jingcai Guo, Deze Zeng, Song Guo
arXiv:2607. 00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecasters, since a proper scoring rule such as the Brier score is computed from outcomes alone and is minimized in expectation by the true probability.
By Sadanand Singh, Allam Reddy, Manan Chopra
arXiv:2504.09192v5 Announce Type: replace
Abstract: The primary goal of my Ph.D. study is to develop provably efficient and practical algorithms for data-driven sequential decision-making under uncer...
By Zhiyong Wang
arXiv:2607. 19518v1 Announce Type: new Abstract: Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search.
By Wouter W. L. Nuijten, Bert de Vries
arXiv:2607. 14407v1 Announce Type: cross Abstract: Many signal processing systems ultimately exist to {act}.
By Osvaldo Simeone