arXiv:2608. 14590v1 Announce Type: new Abstract: LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use of tools.
By Pierre Dantas, Lucas Cordeiro, Ehsan Nowroozi, Tihanyi Norbert
arXiv:2607. 14543v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scenes, follow language instructions, and plan multi-step actions.
By Huaigang Yang, Ya Li, Min Ren, Bo Dai, Zhenliang Zhang, Zhaofeng He
arXiv:2607. 16247v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have empowered embodied agents to execute complex household tasks, they struggle to proactively handle dynamically emerging hazards during closed-loop interactions.
By Bingrui Sima, Lizhong Wang, Xiaoya Lu, Kun He, Xiao Yang
arXiv:2606. 00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions.
By Barak Or
arXiv:2608. 02683v1 Announce Type: cross Abstract: Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accomplish complex tasks.
By Zibo Xiao, Haoyu Wang, Jun Sun
arXiv:2607. 18366v1 Announce Type: new Abstract: Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution.
By Shasha Yu, Fiona Carroll, Barry L. Bentley
CEDAR is a counterexample-guided framework that translates natural-language instructions for embodied agents into regular languages over environment event traces, represented as deterministic finite automata. By using a language model for semantic judgments and execution traces for correction, CEDAR turns constraints into executable finite-state objects, enabling the intersection of learned skills with additional specifications. In Minecraft experiments, CEDAR preserves temporal and spatial constraints better than a program-generating baseline and reduces cumulative LLM queries by reusing learned skills.
By Lekai Chen, Alvaro Velasquez, Ashutosh Trivedi
arXiv:2608. 19729v1 Announce Type: new Abstract: Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety.
By Hyunse Lee, Jiwoo Jeong, Haneul Lee, Kyochul Jang, Youngjae Yu, Woojin Lee
arXiv:2607. 01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks.
By Yunhao Feng, Ruixiao Lin, Ming Wen, Qinqin He, Yanming Guo, Yifan Ding, Yutao Wu, Jialuo Chen, Yunhao Chen, Xiaohu Du, Jianan Ma, Zixing Chen, Zhuoer Xu, Xingjun Ma, Xinhao Deng
arXiv:2606. 23754v1 Announce Type: cross Abstract: Deploying foundation models for robot control raises a central challenge: the expressive power that enables rich, multimodal perception also makes these models opaque and difficult to analyze formally, rendering them intractable for existing verification tools.
By Davide Corsi, Kyungmin Kim, Roy Fox
The paper introduces the Environment State-Text Injection (ESTI) attack, a novel method that manipulates the textual representation of environment states in large language model‑driven embodied agents without altering user instructions, model parameters, or executors. ESTI re‑frames adversarial goals as false state evidence that aligns with the current environment, thereby influencing both planning and execution through object properties, spatial relations, affordances, task‑stage rules, and execution feedback. The authors also present ESTI‑Bench, a benchmark that evaluates attack propagation across the planning‑to‑execution closed loop, and demonstrate that ESTI outperforms existing baselines on multiple embodied task datasets, achieving up to 89.32% higher planning‑level and 43.69% higher execution‑level attack success rates.
By Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu
Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem recently framed as interactive safety. Training such agents to act safely is difficult, since safety and task success are distinct objectives, and safety arises only at a small number of safety-critical steps within a trajectory.