arXiv AI

Shielding for Higher-Order Safety

arXiv:2608. 03662v1 Announce Type: new Abstract: Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety.

arXiv AI
Aug 19

Efficient Dynamic Shielding for Parametric Safety Specifications

The paper presents dynamic shields for AI-controlled autonomous systems, enabling runtime safety enforcement that adapts to changing safety specifications without recomputing from scratch. Unlike traditional static shields, these dynamic shields are pre-designed for a set of possible safety parameters and can quickly adjust as the true specification becomes known during operation. Experiments on robot navigation in unknown terrains show that dynamic shields require only a few minutes offline and a fraction of a second to a few seconds online, outperforming brute-force recomputation by up to five times.

By Davide Corsi, Kaushik Mallik, Andoni Rodriguez, Cesar Sanchez
arXiv AI
Jun 11

Runtime Enforcement of Hybrid System Properties

arXiv:2606. 12022v1 Announce Type: cross Abstract: Runtime enforcement has emerged as a promising approach for ensuring the safety of autonomous and cyber-physical systems operating in uncertain and dynamic environments.

By Mir Md Sajid Sarwar, Srinivas Pinisetty, Rajarshi Ray, Thierry J\'eron
arXiv Machine Learning
Jun 15

Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning

arXiv:2606. 14130v1 Announce Type: new Abstract: Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents.

By Omar Adalat, Edwin Hamel-De le Court, Francesco Belardinelli
arXiv Machine Learning
Sep 3

Exchange Policy Optimization Algorithm for Semi-Infinite Safe Reinforcement Learning

The paper introduces Exchange Policy Optimization (EPO), a framework for semi‑infinite safe reinforcement learning that handles infinitely many constraints by iteratively solving finite subproblems. EPO expands or deletes constraints based on tolerance violations and Lagrange multipliers, maintaining computational tractability while converging to an optimal policy with bounded safety violations. The authors prove finite convergence, provide iteration bounds, and quantify the suboptimality gap under mild assumptions.

By Jiaming Zhang, Yujie Yang, Haoning Wang, Liping Zhang, Shengbo Eben Li