The paper presents dynamic shields for AI-controlled autonomous systems, enabling runtime safety enforcement that adapts to changing safety specifications without recomputing from scratch. Unlike traditional static shields, these dynamic shields are pre-designed for a set of possible safety parameters and can quickly adjust as the true specification becomes known during operation. Experiments on robot navigation in unknown terrains show that dynamic shields require only a few minutes offline and a fraction of a second to a few seconds online, outperforming brute-force recomputation by up to five times.
By Davide Corsi, Kaushik Mallik, Andoni Rodriguez, Cesar Sanchez
arXiv:2606. 13621v1 Announce Type: new Abstract: Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions.
By Achraf Hsain, Sultan Almuhammadi
Shielded reinforcement learning is typically presented as a runtime safety mechanism that compiles temporal-logic specifications into automata restricting an agent's actions. We argue this is the wrong product.
arXiv:2606. 12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dynamic environments.
By Mir Md Sajid Sarwar, Srinivas Pinisetty, Rajarshi Ray, Thierry J\'eron
arXiv:2607. 15003v1 Announce Type: new Abstract: The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.
By Riccardo Curcio, Toni Mancini, Enrico Tronci
arXiv:2511.02605v3 Announce Type: replace
Abstract: Shielding is widely used to enforce safety in reinforcement learning (RL), ensuring that an agent's actions remain compliant with formal specificat...
By Tiberiu-Andrei Georgescu, Alexander W. Goodall, Dalal Alrajeh, Francesco Belardinelli, Sebastian Uchitel
arXiv:2606. 14130v1 Announce Type: new Abstract: Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents.
By Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli
arXiv:2603. 15282v2 Announce Type: replace Abstract: Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees.
By Johannes Schmalz, Chaahat Jain
arXiv:2606. 28639v2 Announce Type: replace-cross Abstract: We establish the mathematical limits of AGI safety in two forms: verifying a fixed system, and verifying that a certified safety property persists once the system self-modifies.
By Jose Pascual Gumbau Mezquita
arXiv:2606. 03804v1 Announce Type: new Abstract: Safe exploration is a key challenge in Reinforcement Learning (RL) that aims to prevent agents from making harmful decisions while exploring their environment.
By Stefan Pranger, Bettina K\"onighofer
The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.
By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li
arXiv:2609.38016v1 Announce Type: new
Abstract: Constrained Reinforcement Learning has recently gained increasing attention in the field of Safe Autonomous Driving, where the general mechanism is to...
By Huan Rong, Chao Yin, Anouar Imel, Yijie Xia, Tinghuai Ma