arXiv Machine Learning

Deep Reinforcement Learning for Reach-Avoid-Stay Problems

The paper introduces a two‑step deep reinforcement learning framework for Reach‑Avoid‑Stay (RAS) problems, aiming to compute the maximal robust RAS set and its control policy for general dynamic systems. First, it learns the maximal robust control‑invariant set inside the target and a policy to keep the system within it. Then it uses this invariant set as a target to compute the maximal robust reach‑avoid set, proving equivalence to the maximal robust RAS set and constructing a switching policy that guarantees task completion. Simulation results show the method achieves exact maximal RAS sets without training errors and outperforms baseline approaches in accuracy and performance.

arXiv Machine Learning
5d ago

Robust Successor Features

The paper introduces robust successor features, a method that extends the successor representation to handle uncertainty in both reward functions and transition kernels within linear Markov Decision Processes. It provides a theoretical bound on Generalized Policy Improvement that quantifies performance loss due to mismatched dynamics, and demonstrates the approach on grid-based benchmarks against prior methods that consider only reward or transition differences.

By Erik Nikulski, Yamen Habib, Vicen\c{c} Gomez, Anders Jonsson, Rub\'en Moreno-Bote, Javier Segovia-Aguas
arXiv Machine Learning
Sep 24

LEAP-CBF: A Safety Filter for Uncertain Systems with Least-Effort Adversarial Potentials

The paper introduces LEAP-CBF, a safety filter that uses Least‑Effort Adversarial Potentials to quantify how much disturbance effort is needed to cause failure in nonlinear dynamical systems. LEAP serves as a control barrier function for the undisturbed system and can be combined with a robust safety filter that tolerates disturbances with bounded cumulative effort. The authors develop a deep reinforcement learning method to construct LEAPs and validate their effectiveness through simulations of multi‑agent systems and hardware experiments on a quadruped and quadrotors.

By Oswin So, Eric Yu, Chuchu Fan
arXiv Machine Learning
Aug 28

Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

The paper introduces Safe Contrastive Reinforcement Learning (Safe-CRL), a method that corrects bias in contrastive RL caused by failure-terminated Markov decision processes. By applying mass-weighted InfoNCE and a log-survival-mass score, Safe-CRL uses only a one-bit failure signal to improve survival and goal-reaching performance across twelve robot navigation and locomotion tasks. The approach demonstrates complex failure-avoidance behaviors and completes the theoretical foundation of contrastive RL under failure termination.

By Guopeng Li, Yiyang Duan, Yiru Jiao, Chengcheng Xu
arXiv Machine Learning
Sep 14

A Unified and Constrained View of Regularization-Based Robust Reinforcement Learning

The paper presents a unified framework for regularization-based robust reinforcement learning by deriving upper bounds on the performance gap between nominal and worst-case policies. These bounds are expressed as a regularization objective plus a KL-divergence penalty, explaining why KL penalties enhance robustness. The authors reformulate robust training as a constrained optimization problem, updating the Lagrange multiplier jointly with the policy to automatically tune regularization, and validate the approach with extensive adversarial evaluations on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou
arXiv Machine Learning
Sep 14

Robust Policy Optimization via Adversarial Importance Sampling

The paper introduces Adversarial Importance Sampling (Advis), a technique that leverages importance sampling over standard training trajectories to estimate and optimize worst‑case returns without extra environment interactions or auxiliary networks, thereby capturing long‑term robustness. It also presents advrl, a modular PyTorch library that consolidates existing robustness methods and adversarial attacks into single‑file implementations for easier prototyping and reproducible evaluation. Finally, the authors highlight that optimal adversarial hyperparameters do not transfer across agents, prompting evaluation against a broader set of attackers (6–14× more configurations) and demonstrate the effectiveness of their approach on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou