Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning
arXiv:2609. 12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition.
arXiv:2607. 13175v1 Announce Type: cross Abstract: Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events.
arXiv:2609. 12014v1 Announce Type: new Abstract: Safe offline reinforcement learning assumes a cost function on every transition.
The paper introduces new evaluation metrics for safe reinforcement learning that go beyond average safety guarantees by examining how often and how severely safety bounds are violated, consistency across tasks and bounds, and the relationship between training-time and final policy behavior. It also proposes a safety tier system for categorizing algorithms and presents empirical safety evaluations on multiple navigation tasks. The authors recommend reporting aggregate metrics, distributional data, and task‑specific results together, and provide an open‑source suite, SafeRLEval, to facilitate reliable safety assessment.
arXiv:2603. 10938v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation captures only a single statistic of the cost distribution and fails to account for distributional uncertainty, particularly under heavy tails or rare catastrophic events.
BCPPO is a new variant of Proximal Policy Optimization that uses Bachelier-inspired cost‑prediction networks to generate a smooth penalty based on disagreement among critics. The method keeps temporal‑difference learning unchanged, applies a saturation‑aware controller to manage cost penalties, and deploys only the policy network. Across extensive experiments, BCPPO outperforms comparators in achieving higher mean returns while maintaining lower or comparable CVaR in all tested tasks.
arXiv:2603. 15136v2 Announce Type: replace-cross Abstract: Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints.
arXiv:2606. 10228v1 Announce Type: cross Abstract: Safe exploration is a prerequisite for deploying reinforcement learning (RL) agents in safety-critical domains.
arXiv:2606. 31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics.
The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.
arXiv:2606. 04812v1 Announce Type: cross Abstract: Guaranteeing safety is critical to the deployment of reinforcement learning (RL) agents in the real-world, especially as policies learned using deep RL may demonstrate susceptibility to transition perturbations that result in unknown or unsafe behaviour.
arXiv:2601. 22993v4 Announce Type: replace Abstract: We introduce Canary, a risk-averse method designed to optimize Value-at-Risk (VaR) constrained reinforcement learning (RL) problems.
arXiv:2607. 07976v1 Announce Type: cross Abstract: Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs).
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.