arXiv AI

Interval POMDP Shielding for Imperfect-Perception Agents

The paper introduces Interval POMDP Shielding for agents with imperfect perception, aiming to prevent unsafe actions when sensor readings may be misclassified. By estimating perception uncertainty from finite labeled data, the authors construct confidence intervals and model the system as a finite Interval Partially Observable Markov Decision Process. They propose an algorithm that computes a conservative belief set, enabling a runtime shield that guarantees, with high probability, that any action allowed by the shield meets a specified safety lower bound. Experiments on four case studies demonstrate that this shielding approach outperforms state‑of‑the‑art baselines in safety.

arXiv AI
Aug 21

Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning

arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).

By Astrid Horn Brorholt (Aalborg University, Aalborg, Denmark), Maris F. L. Galesloot (Radboud University, Nijmegen, Netherlands), Nils Jansen (Radboud University, Nijmegen, Netherlands), Kim Guldstrand Larsen (Aalborg University, Aalborg, Denmark), Christian Schilling (Aalborg University, Aalborg, Denmark)
arXiv Machine Learning
Sep 22

Correcting Learning-based Perception for Safety

The paper presents a two-step method to correct machine‑learning based perception for safety in autonomous systems. First, it uses offline computation to characterize uncertainties from the ML module via preimages of perception contracts. Then, at runtime, a risk heuristic selects specific states from these uncertain estimates to guide control decisions, reducing safety violations in adaptive cruise control scenarios while adding minimal delay.

By Yan Miao, Hussein Darir, Sayan Mitra
arXiv Machine Learning
Jun 2

All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning

arXiv:2606. 01363v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) infers information about the environment from a learned dynamics model and bears the potential to address open problems such as data efficient and safe learning in robotics.

By Bernd Frauenknecht, Devdutt Subhasish, Artur Eisele, Friedrich Solowjow, Sebastian Trimpe