arXiv AI

Safe and Robust Neural Policy Learning with Statistical Verification for Sim-to-Real Deployment in Robotics

Hugging Face Trending Papers
Aug 19

SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation

SCAPE is a scenario‑conditioned simulation‑augmented policy evaluation framework that predicts real‑world policy performance for specific scenarios using limited paired simulation‑and‑real samples and extensive simulation rollouts. It corrects sim‑to‑real bias in simulation labels before training the prediction model and calibrates prediction uncertainty via conformal prediction. Experiments on autonomous driving and quadruped velocity tracking show SCAPE reduces scenario‑level prediction error, improves testing sample efficiency, narrows calibrated prediction intervals, and generalizes better to out‑of‑distribution scenarios, enabling fine‑grained deployment strategies.

Hugging Face Trending Papers
Jun 17

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks.

arXiv AI
Sep 24

Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning

The paper introduces Safety to Competence (S2C), a two‑stage reinforcement learning framework that first learns a safety filter and then trains a competitive task policy while embedding the filter. By separating safety synthesis from task learning, S2C reduces training complexity and prevents the policy from being exploited by adversarial attacks. Experiments on simulated touchdown games show that S2C achieves higher win rates, better Elo ratings, and lower exploitability than eight safe‑RL baselines, and hardware tests confirm its competence against a human opponent.

By Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu
arXiv AI
Jun 19

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

arXiv:2606. 19632v1 Announce Type: cross Abstract: Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets.

By Ahmad Farooq, Kamran Iqbal
arXiv AI
Sep 17

CALOS: Control-Affine Lyapunov On-manifold Safety Layer for Safe Deep Reinforcement Learning for Quadrotors

CALOS is a runtime safety layer for quadrotor control that enforces attitude constraints without altering the underlying deep reinforcement learning algorithm. It formulates tilt-angle inequalities and a Lyapunov descent condition into a single quadratic program, solved exactly via active-set enumeration over a three-dimensional torque space. In NVIDIA Isaac Lab trajectory-tracking tasks, CALOS reduces lateral tracking error by 55‑60% compared to an unconstrained Proximal Policy Optimization baseline and eliminates attitude-constraint violations during training.

By Fabrizio Cesareo, Sebastiano Mengozzi, Nicola Mimmo, Andrea Acquaviva