Robust Shielding for Safe Reinforcement Learning
arXiv:2606. 00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs).
The paper introduces Interval POMDP Shielding for agents with imperfect perception, aiming to prevent unsafe actions when sensor readings may be misclassified. By estimating perception uncertainty from finite labeled data, the authors construct confidence intervals and model the system as a finite Interval Partially Observable Markov Decision Process. They propose an algorithm that computes a conservative belief set, enabling a runtime shield that guarantees, with high probability, that any action allowed by the shield meets a specified safety lower bound. Experiments on four case studies demonstrate that this shielding approach outperforms state‑of‑the‑art baselines in safety.
arXiv:2606. 00270v1 Announce Type: new Abstract: Shielding is an effective approach to formally guarantee the safety of reinforcement learning agents in Markov decision processes (MDPs).
arXiv:2608. 19836v1 Announce Type: cross Abstract: Probabilistic shielding is a technique for safe reinforcement learning (RL).
The paper presents a two-step method to correct machine‑learning based perception for safety in autonomous systems. First, it uses offline computation to characterize uncertainties from the ML module via preimages of perception contracts. Then, at runtime, a risk heuristic selects specific states from these uncertain estimates to guide control decisions, reducing safety violations in adaptive cruise control scenarios while adding minimal delay.
arXiv:2602. 23545v2 Announce Type: replace Abstract: In the real world, planning is often challenged by distribution shifts.
arXiv:2606. 04634v1 Announce Type: new Abstract: Trust in a decision-making system requires both safety guarantees and the ability to interpret and understand its behavior.
arXiv:2609.35874v1 Announce Type: new Abstract: Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost state...
arXiv:2609.15915v1 Announce Type: new Abstract: Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL...
arXiv:2606. 02562v1 Announce Type: cross Abstract: Autonomous robots that interact with people must make safe and efficient decisions under human-induced uncertainty, such as their preferences, goals, competency, and willingness to cooperate.
arXiv:2606. 01363v1 Announce Type: new Abstract: Model-based reinforcement learning (MBRL) infers information about the environment from a learned dynamics model and bears the potential to address open problems such as data efficient and safe learning in robotics.
arXiv:2307. 10524v3 Announce Type: replace Abstract: We study the tradeoff between consistency and robustness in the context of a single-trajectory time-varying Markov Decision Process (MDP) with untrusted machine-learned advice.
arXiv:2601. 18930v4 Announce Type: replace-cross Abstract: We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms.
arXiv:2607. 04019v1 Announce Type: cross Abstract: Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system can reliably use.