arXiv:2606. 16605v1 Announce Type: new Abstract: World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making.
By Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu
TrojanWorld is a backdoor framework that targets world-model agents by steering their internal imagination toward attacker-specified actions when a physical trigger is present. The attack uses Decision-Reflective Induction, Clean Behavior Anchoring, and Causal Propagation to maintain stealth, persistence, and high performance. Experiments on TD-MPC2, DreamerV3, and R2-Dreamer across several benchmarks show that the attack can induce target actions with minimal performance loss and can keep agents on a malicious trajectory even after the trigger is removed.
By Wenkai Huang, Siyuan Liang, Gaolei Li, Yiming Li, Tianhao Peng, Jianhua Li, Dacheng Tao
arXiv:2607. 15207v1 Announce Type: new Abstract: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction.
By Qi Li, Xingyi Yang, Xinchao Wang
The paper introduces the Environment State-Text Injection (ESTI) attack, a novel method that manipulates the textual representation of environment states in large language model‑driven embodied agents without altering user instructions, model parameters, or executors. ESTI re‑frames adversarial goals as false state evidence that aligns with the current environment, thereby influencing both planning and execution through object properties, spatial relations, affordances, task‑stage rules, and execution feedback. The authors also present ESTI‑Bench, a benchmark that evaluates attack propagation across the planning‑to‑execution closed loop, and demonstrate that ESTI outperforms existing baselines on multiple embodied task datasets, achieving up to 89.32% higher planning‑level and 43.69% higher execution‑level attack success rates.
By Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu
arXiv:2606. 18697v1 Announce Type: new Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments.
By Yibin Hu, Xiaolin Sun, Zizhan Zheng
The article discusses how current world models, while achieving high predictive likelihood and visual fidelity, often fail to preserve the evidence needed for safe decision-making in embodied systems. It identifies three structural mismatches—likelihood versus risk, prediction versus intervention, and finite-horizon prediction versus accumulated consequences—and proposes the Risk‑Informed World Model (RIWM) as a decision‑centric framework. RIWM emphasizes consequences, intervention, epistemic uncertainty, and recoverability, integrating decision‑relevant representation, counterfactual reasoning, safety‑critical episodic memory, and runtime safety assurance to better support safety‑critical embodied systems.
By Kailang Ma, Heye Huang, Inhi Kim, Kitae Jang
arXiv:2606. 09499v1 Announce Type: cross Abstract: World models have recently seen a rapid growth in both their popularity and capability as more data efficient tools for generating robot training data or simulating real world environments, with many works proposing their integration into the robot learning pipeline.
By Ethan Rathbun, Ahmed Agha, Saaduddin Mahmud, Christopher Amato, Alina Oprea, Eugene Bagdasarian
arXiv:2608. 05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services.
By Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu
The paper introduces LogiC-Diff, a logic-conditioned bi-stage diffusion framework that embeds Signal Temporal Logic (STL) specifications into AI-enabled cyber‑physical system (CPS) forecasting models. By using STL as a conditioning signal, the method repairs inputs and refines outputs to jointly mitigate adversarial perturbations and enforce desired temporal behaviors. Experiments on two real‑world CPS datasets show that LogiC-Diff consistently improves robustness and specification compliance across various sensor faults and cyber attacks, outperforming reconstruction‑based defenses.
By Ziyan An, John Stankovic, Meiyi Ma
arXiv:2605. 31119v2 Announce Type: replace-cross Abstract: In robotics, dangers and adversity modes are often embodiment-specific and relative to each agent.
By Navin Sriram Ravie, Andrew Jong, Krrish Jain, John Liu, Omar Alama, Bijo Sebastian, Sebastian Scherer
arXiv:2606. 01991v1 Announce Type: new Abstract: As Large Language Model (LLM) agents increasingly leverage the Model Context Protocol (MCP) to operate in complex environments, the expansion of their action spaces offers agents unsafe capabilities and underscores the risk of power-seeking.
By Lichao Wang, Zhaoxing Ren, Tianzhuo Yang, Jiaming Ji, Chi Harold Liu, Yaodong Yang, Juntao Dai
Future-Back Threat Modeling (FBTM) is a predictive security framework that starts with envisioned future threat states and works backward to uncover assumptions, gaps, blind spots, and vulnerabilities in current defense architectures. It aims to reveal both known unknowns and unknown unknowns, including emerging tactics, techniques, and procedures, thereby improving the predictability of adversary behavior under future uncertainty. By anticipating future threats such as AI, information warfare, and supply chain attacks, FBTM helps security leaders make informed decisions today to build more resilient security postures for the future.
By Vu Van Than