arXiv Machine Learning By Mehrdad Moghimi, Bernardo Avila Pires

Utility-Constrained Policy Optimization

Read the original on arXiv Machine Learning →

arXiv:2606. 14029v1 Announce Type: new Abstract: Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents; however, the framework does not support risk-sensitive constraints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 27

Safe In-Context Reinforcement Learning

arXiv:2509. 25582v4 Announce Type: replace Abstract: In-context reinforcement learning (ICRL) is an emerging RL paradigm where an agent, after pretraining, can adapt to out-of-distribution test tasks without any parameter updates, instead relying on an expanding context of interaction history.

By Amir Moeini, Minjae Kwon, Alper Kamil Bozkurt, Yuichi Motai, Rohan Chandra, Lu Feng, Shangtong Zhang
arXiv Machine Learning
Sep 3

Cantelli Constrained Policy Optimization

The paper introduces Canary, a risk‑averse reinforcement learning method that optimizes Value‑at‑Risk (VaR) constraints. By applying Cantelli’s inequality, Canary derives a tractable, conservative, and smooth bound on the VaR constraint using only the first two moments of the cost return, yielding a stable constraint estimator even with tight violation thresholds. Extending the trust‑region framework of Constrained Policy Optimization (CPO), the authors provide worst‑case bounds for policy improvement and constraint violation, and empirically demonstrate that Canary reliably satisfies the VaR constraint in every tested environment.

By Rohan Tangri, Jan-Peter Calliess
arXiv AI
Sep 15

Evaluation Metrics for Safe Reinforcement Learning

The paper introduces new evaluation metrics for safe reinforcement learning that go beyond average safety guarantees by examining how often and how severely safety bounds are violated, consistency across tasks and bounds, and the relationship between training-time and final policy behavior. It also proposes a safety tier system for categorizing algorithms and presents empirical safety evaluations on multiple navigation tasks. The authors recommend reporting aggregate metrics, distributional data, and task‑specific results together, and provide an open‑source suite, SafeRLEval, to facilitate reliable safety assessment.

By Lindsay Spoor, Aske Plaat, Thomas Moerland