arXiv Machine Learning

Verifiable Foundation Models for Robot Safety

arXiv:2606. 23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes these models opaque and difficult to analyze formally, rendering them intractable for existing verification tools.

arXiv AI
Sep 12

ActSafeGuard: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies

arXiv:2609. 11697v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment.

By Jianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang, Yiheng Li, Yue Gao
arXiv AI
Sep 25

CrossSafe: Towards Cross-Embodiment Latent Safety Filters

CrossSafe proposes embodiment-conditioned safety filtering that uses a Hamilton‑Jacobi reachability value function shared across robots while conditioning on each robot’s morphology and kinematics via a morphology‑aware latent representation. The method performs reachability analysis directly in latent space, enabling a single policy trained on multiple bimanual robot embodiments and manipulation tasks to generalize zero‑shot to a held‑out embodiment and reduce collision rates. Experiments on five embodiments and five tasks demonstrate that training with more embodiments improves generalization.

By Ihab Tabbara, Yuxuan Yang, Hussein Sibai
arXiv AI
Sep 24

Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning

The paper introduces Safety to Competence (S2C), a two‑stage reinforcement learning framework that first learns a safety filter and then trains a competitive task policy while embedding the filter. By separating safety synthesis from task learning, S2C reduces training complexity and prevents the policy from being exploited by adversarial attacks. Experiments on simulated touchdown games show that S2C achieves higher win rates, better Elo ratings, and lower exploitability than eight safe‑RL baselines, and hardware tests confirm its competence against a human opponent.

By Ruihan Wu, Rui Yang, Donggeon David Oh, Duy Nguyen, Haimin Hu
arXiv AI
Jul 17

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

arXiv:2607. 14543v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions.

By Huaigang Yang, Ya Li, Min Ren, Bo Dai, Zhenliang Zhang, Zhaofeng He
arXiv AI
Jul 28

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

arXiv:2607. 23784v1 Announce Type: cross Abstract: While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail.

By Daphne Chen, Archit Ritesh Jain, Eric Goossen, Emma Romig, Michael Murray, Nick Walker, Maya Cakmak
Hugging Face Trending Papers
Jun 17

Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation

Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets. We present the first end-to-end framework for safety verification of learned multi-agent communication policies through policy abstraction: neural policies are distilled into interpretable decision trees, then formally verified, with empirical validation confirming that verified safety properties transfer to original networks.