arXiv:2602. 03778v2 Announce Type: replace-cross Abstract: Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catastrophic events.
By Aneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon, Erick Delage
arXiv:2601. 22993v4 Announce Type: replace Abstract: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
By Rohan Tangri, Jan-Peter Calliess
The paper presents a policy gradient theorem tailored to Cumulative Prospect Theory (CPT) objectives in finite-horizon reinforcement learning, extending the classic policy gradient framework to include distortion-based risk measures. Leveraging this theorem, the authors develop a first‑order policy gradient algorithm that uses a Monte Carlo estimator based on order statistics, providing statistical guarantees and proving asymptotic convergence to first‑order stationary points of the generally nonconvex CPT objective. The work also offers a non‑asymptotic sample complexity bound for reaching an approximate stationary policy and demonstrates the qualitative effects of CPT through simulations, comparing the new first‑order method to existing zeroth‑order approaches.
By Olivier Lepel, Anas Barakat
The paper introduces Canary, a risk‑averse reinforcement learning method that optimizes Value‑at‑Risk (VaR) constraints. By applying Cantelli’s inequality, Canary derives a tractable, conservative, and smooth bound on the VaR constraint using only the first two moments of the cost return, yielding a stable constraint estimator even with tight violation thresholds. Extending the trust‑region framework of Constrained Policy Optimization (CPO), the authors provide worst‑case bounds for policy improvement and constraint violation, and empirically demonstrate that Canary reliably satisfies the VaR constraint in every tested environment.
By Rohan Tangri, Jan-Peter Calliess
arXiv:2609.24103v1 Announce Type: new
Abstract: In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. W...
By Larry Preuett, Qiuyi Zhang, Muhammad Aurangzeb Ahmad
arXiv:2608. 03562v1 Announce Type: new Abstract: Reinforcement learning (RL) with general utility extends classic RL by optimizing an arbitrary utility functional of the policy-induced occupancy measure, thereby enabling a broader range of applications.
By Zixuan Liu, Fangzheng Wu, Brian Summa, Zizhan Zheng
arXiv:2609. 10866v1 Announce Type: new Abstract: Reinforcement learning (RL) agents deployed in real-world environments are often vulnerable to adversarial perturbations in state observations, creating risks in safety-critical applications.
By Tong Li, Saunak Kumar Panda, Yisha Xiang
arXiv:2609.38938v1 Announce Type: new
Abstract: Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies...
By Xinyi Ni, Lifeng Lai
In many real-world planning tasks, agents must tackle uncertainty about the environment's state and variability in the outcomes of any chosen policy. We address both forms of uncertainty as a first st...
arXiv:2607. 09298v1 Announce Type: cross Abstract: We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives.
By Pedro P. Santos, F\'abio Vital, Alberto Sardinha, Francisco S. Melo
arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.
By Yaswanth Chittepu, Ativ Joshi, Rajarshi Bhattacharjee, Scott Niekum
arXiv:2607. 14373v1 Announce Type: new Abstract: We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics.
By Yang Liu, Yuhao Liu, Yunran Wei