We’re releasing the public beta of OpenAI Gym, a toolkit for developing and comparing reinforcement learning (RL) algorithms. It consists of a growing suite of environments (from simulated robots to Atari games), and a site for comparing and reproducing results.
The paper introduces new evaluation metrics for safe reinforcement learning that go beyond average safety guarantees by examining how often and how severely safety bounds are violated, consistency across tasks and bounds, and the relationship between training-time and final policy behavior. It also proposes a safety tier system for categorizing algorithms and presents empirical safety evaluations on multiple navigation tasks. The authors recommend reporting aggregate metrics, distributional data, and task‑specific results together, and provide an open‑source suite, SafeRLEval, to facilitate reliable safety assessment.
By Lindsay Spoor, Aske Plaat, Thomas Moerland
arXiv:2606. 28347v1 Announce Type: cross Abstract: Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming.
By Charles L. Wang, Keir Dorchen, Peter Jin
RePolicy is a reinforcement learning approach designed to invoke safety policies for language model agents by evaluating entire execution trajectories within context-dependent policy libraries. It generates policy-grounded rationales and safety judgments, and is initialized with the PolicyTraj-20K dataset before fine-tuning via GRPO with verifiable rewards and policy-context perturbation. Experiments on six safety benchmarks demonstrate strong safety-detection performance and robust policy invocation across varying contexts.
By Houcheng Jiang, Boxuan Zhang, Qiyong Zhong, Junfeng Fang, Xiang Wang, Xiangnan He
arXiv:2609.15915v1 Announce Type: new
Abstract: Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL...
By Zeyang Li, Sunbochen Tang, Navid Azizan
arXiv:2606. 31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics.
By Hongpeng Cao, Liqun Zhao, Yuliang Gu, Naira Hovakimyan, Lui Sha, Marco Caccamo
The paper proves that using a permissive safety filter in reinforcement learning does not compromise asymptotic performance. By formalizing safety through a safety‑critical Markov decision process and a filtered MDP, the authors show that optimal policies in the filtered MDP achieve the same return as the best safe policy in the original setting. Experiments on Safety Gymnasium confirm zero violations during training and performance that matches or exceeds unfiltered baselines.
By Donggeon David Oh, Duy P. Nguyen, Haimin Hu, Jaime Fern\'andez Fisac
Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised f...
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
By Kaustubh Mani, Yann Pequignot, Vincent Mai, Liam Paull
RL-Teacher is an open-source implementation of our interface to train AIs via occasional human feedback rather than hand-crafted reward functions. The underlying technique was developed as a step towards safe AI systems, but also applies to reinforcement learning problems with rewards that are hard to specify.
arXiv:2607. 07029v1 Announce Type: cross Abstract: Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks.
By Dennis Gross, Quentin Mazouni, Helge Spieker, Arnaud Gotlieb
The paper introduces Safety to Competence (S2C), a two‑stage reinforcement learning framework that first learns a safety filter and then trains a competitive task policy while embedding the filter. By separating safety synthesis from task learning, S2C reduces training complexity and prevents the policy from being exploited by adversarial attacks. Experiments on simulated touchdown games show that S2C achieves higher win rates, better Elo ratings, and lower exploitability than eight safe‑RL baselines, and hardware tests confirm its competence against a human opponent.
By Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu