arXiv AI

SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

arXiv:2510. 12985v3 Announce Type: replace Abstract: We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents.

arXiv AI
Jul 17

SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents

arXiv:2607. 14543v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions.

By Huaigang Yang, Ya Li, Min Ren, Bo Dai, Zhenliang Zhang, Zhaofeng He
arXiv Computation and Language
Aug 31

CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action

CEDAR is a counterexample-guided framework that translates natural-language instructions for embodied agents into regular languages over environment event traces, represented as deterministic finite automata. By using a language model for semantic judgments and execution traces for correction, CEDAR turns constraints into executable finite-state objects, enabling the intersection of learned skills with additional specifications. In Minecraft experiments, CEDAR preserves temporal and spatial constraints better than a program-generating baseline and reduces cumulative LLM queries by reusing learned skills.

By Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi
arXiv Machine Learning
Jun 24

Verifiable Foundation Models for Robot Safety

arXiv:2606. 23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes these models opaque and difficult to analyze formally, rendering them intractable for existing verification tools.

By Davide Corsi, Kyungmin Kim, Roy Fox
arXiv AI
Aug 20

Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents

The paper introduces the Environment State-Text Injection (ESTI) attack, a novel method that manipulates the textual representation of environment states in large language model‑driven embodied agents without altering user instructions, model parameters, or executors. ESTI re‑frames adversarial goals as false state evidence that aligns with the current environment, thereby influencing both planning and execution through object properties, spatial relations, affordances, task‑stage rules, and execution feedback. The authors also present ESTI‑Bench, a benchmark that evaluates attack propagation across the planning‑to‑execution closed loop, and demonstrate that ESTI outperforms existing baselines on multiple embodied task datasets, achieving up to 89.32% higher planning‑level and 43.69% higher execution‑level attack success rates.

By Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu
Hugging Face Trending Papers
Aug 20

SafeBranch: Branch-Pair Safety Alignment for Embodied Agents

Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are distinct objectives, and safety arises only at a small number of safety-critical steps within a trajectory.