arXiv:2605. 27628v2 Announce Type: replace Abstract: As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge.
By Srini Ramaswamy
arXiv:2606. 00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions.
By Barak Or
arXiv:2607. 26121v1 Announce Type: cross Abstract: Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction.
By Xinyu Yang, Tianxing Chen, Honghao Su, Minxuan Wang, Chenze Yu, Zhangzheng Tu, Yue Chen, Yuxiao Huo, Lingfeng Zhang, Yan Huang, Yan Qin, Shaolong Zhu, Qiwei Liang, Hekun Tian, Shujia Liu, Guangyu Chen, Junhao Gong, Zixuan Li, Wenwei Lin, Zijian Lin, Wenxuan Zhu, Eric J Chen, Yue Yuan, Qize Yu, Jiaqi Liang, Haowen Yan, Hengfei Zhao, Weijie Wan, Zikun Xiao, Junyuan Tang, Baijun Chen, Kai-Chong Lei, Kaixuan Wang, Kailun Su, Zanxin Chen, Yao Mu, Renjing Xu, Chuqiao Lyu, Qi Xiong, Ping Luo, Wenbo Ding
SafeEvolve is an experience-driven framework that co‑evolves a harness and policy to improve safety alignment for LLM‑based agents. It uses on‑policy trajectory safety evidence to update safety prompts and hierarchical skills, producing auditable harness artifacts. The policy is trained via a two‑stage SFT‑RL pipeline that bootstraps with the evolved harness and then refines behavior through verifier‑decomposed rewards, yielding a better safety‑utility tradeoff on benchmarks such as AgentDojo.
arXiv:2606. 26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems.
By Seth Dobrin, {\L}ukasz Chmiel
SafeEvolve is an experience-driven framework that co‑evolves a harness and policy to align large‑language‑model agents with safety goals. It uses completed on‑policy trajectories to update safety prompts and hierarchical skills, then applies a two‑stage SFT‑RL training loop that bootstraps the policy with the evolved harness and refines it through verifier‑augmented rewards. Experiments on agentic safety benchmarks show that SafeEvolve improves the safety‑utility tradeoff, achieving a three‑fold reduction in ASR on AgentDojo for Qwen3.5‑4B while increasing benign utility from 59.79% to 61.86%.
By Qinghua Mao, Wanying Qu, Dadi Guo, Leitao Yuan, Qingyu Liu, Yu Li, Guanxu Chen, Yanwei Fu, Xi Lin, Xia Hu, Dongrui Liu
arXiv:2607. 25408v1 Announce Type: new Abstract: A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.
By Debjyoti Paul
arXiv:2606. 05660v1 Announce Type: cross Abstract: Embodied AI systems are increasingly expected to reason and act over extended horizons in physical environments.
By Dabin Kim, Daemin Park, Sangyub Lee, Jinsik Kim, Yeongtak Oh, Jongho Shin, Sungroh Yoon
arXiv:2606. 15563v1 Announce Type: new Abstract: AI systems increasingly delegate decisions to specialized models, evaluators, tools, and supervisory controllers.
By Carlos R. B. Azevedo
arXiv:2606. 02641v1 Announce Type: cross Abstract: Interactive driving exposes a failure mode that is easy to miss in rule-aware autonomous-driving stacks: a hard-rule margin can be negative for an ego candidate even though a small lawful accommodation by a non-priority agent would restore feasibility.
By Yifan Wang
The paper introduces a finite‑sample probabilistic safety certification framework for black‑box AI decision models used in closed‑loop grid operation. It transforms the AI‑grid evaluation into a binary unsafe outcome under a safety specification and applies exact binomial inference to provide a tight one‑sided upper bound on the unsafe operation probability, using held‑out calibration scenarios. The framework also incorporates physically interpretable sample‑space adversarial attacks to address distribution shifts and is validated through case studies involving 1,000‑agent AI models for grid‑edge flexibility coordination.
By Yihong Zhou, Hanbin Yang, Thomas Morstyn
arXiv:2605. 17909v2 Announce Type: replace Abstract: As autonomous agentic systems scale across regulated critical infrastructures, the lack of mechanistic, hardware-rooted enforcement for high-frequency policy updates presents a fundamental safety gap.
By Riddhi Mohan Sharma