Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. However, since learning-based control may not be able to realize safety-guaranties, it is of great importance to enhance safety and robustness while maintaining good performances.
arXiv:2607. 07252v1 Announce Type: new Abstract: Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics.
By Georg Sch\"afer, Jakob Rehrl, Stefan Huber, Simon Hirlaender
Reinforcement learning (RL) enables the synthesis of control policies directly from data, making it highly appealing for complex cyber-physical systems (CPSs) and robotics. A persistent challenge, however, is ensuring strict, hard safety constraints during the active learning phase.
arXiv:2606. 31320v1 Announce Type: new Abstract: Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics.
By Hongpeng Cao, Liqun Zhao, Yuliang Gu, Naira Hovakimyan, Lui Sha, Marco Caccamo
CALOS is a runtime safety layer for quadrotor control that enforces attitude constraints without altering the underlying deep reinforcement learning algorithm. It formulates tilt-angle inequalities and a Lyapunov descent condition into a single quadratic program, solved exactly via active-set enumeration over a three-dimensional torque space. In NVIDIA Isaac Lab trajectory-tracking tasks, CALOS reduces lateral tracking error by 55‑60% compared to an unconstrained Proximal Policy Optimization baseline and eliminates attitude-constraint violations during training.
By Fabrizio Cesareo, Sebastiano Mengozzi, Nicola Mimmo, Andrea Acquaviva
arXiv:2605. 26452v2 Announce Type: replace-cross Abstract: Safe reinforcement learning (RL) for robotic systems requires policies that improve task performance while satisfying state and input constraints during both training and deployment.
By Dhruv S. Kushwaha, Zoleikha A. Biron
arXiv:2408. 12548v3 Announce Type: replace Abstract: Machine Learning (ML) has become central to Autonomous Vehicles (AVs), supporting perception, prediction, planning, control, and decision-making in dynamic environments.
By Yousef Emami, Mohammadhossein Homaei, Miguel Guti\'errez Gait\'an, Luis Almeida, Kai Li, Hui Huang, Zhu Han
arXiv:2501. 15373v2 Announce Type: replace-cross Abstract: Merely pursuing performance may adversely affect safety, while a conservative policy for safe exploration will degrade the performance.
By Xinyang Wang, Hongwei Zhang, Shimin Wang, Wei Xiao, Martin Guay
arXiv:2606. 24010v1 Announce Type: new Abstract: Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints.
By Zihao Guo, Jianing Zhao, Ling Li, Hao Liang, Giuseppe Loianno, Yali Du
The paper presents a two-step method to correct machine‑learning based perception for safety in autonomous systems. First, it uses offline computation to characterize uncertainties from the ML module via preimages of perception contracts. Then, at runtime, a risk heuristic selects specific states from these uncertain estimates to guide control decisions, reducing safety violations in adaptive cruise control scenarios while adding minimal delay.
By Yan Miao, Hussein Darir, Sayan Mitra
The paper introduces LEAP-CBF, a safety filter that uses Least‑Effort Adversarial Potentials to quantify how much disturbance effort is needed to cause failure in nonlinear dynamical systems. LEAP serves as a control barrier function for the undisturbed system and can be combined with a robust safety filter that tolerates disturbances with bounded cumulative effort. The authors develop a deep reinforcement learning method to construct LEAPs and validate their effectiveness through simulations of multi‑agent systems and hardware experiments on a quadruped and quadrotors.
By Oswin So, Eric Yu, Chuchu Fan
The paper introduces Safety to Competence (S2C), a two‑stage reinforcement learning framework that first learns a safety filter and then trains a competitive task policy while embedding the filter. By separating safety synthesis from task learning, S2C reduces training complexity and prevents the policy from being exploited by adversarial attacks. Experiments on simulated touchdown games show that S2C achieves higher win rates, better Elo ratings, and lower exploitability than eight safe‑RL baselines, and hardware tests confirm its competence against a human opponent.
By Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu