arXiv AI

ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control

arXiv:2606. 16605v1 Announce Type: new Abstract: World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making.

arXiv AI
Sep 25

Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning

The paper investigates a vulnerability in feedback‑based agent planning, showing that the first round of feedback corrects a large portion of adversarial directions (46%) while subsequent rounds see a sharp decline (13% and 7%). The authors attribute this to an initialization anchoring weakness driven by plausible plan shifts, lack of counterevidence, and persistence of accepted directions. They introduce “InitAnchor”, a black‑box attack framework that exploits these factors, achieving high attack success rates across diverse tasks, architectures, and LLMs, and remaining effective against multiple defenses and real‑world agents.

By Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Peng Zhan, Zheng Li, Shanqing Guo
arXiv Machine Learning
Sep 10

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

TrojanWorld is a backdoor framework that targets world-model agents by steering their internal imagination toward attacker-specified actions when a physical trigger is present. The attack uses Decision-Reflective Induction, Clean Behavior Anchoring, and Causal Propagation to maintain stealth, persistence, and high performance. Experiments on TD-MPC2, DreamerV3, and R2-Dreamer across several benchmarks show that the attack can induce target actions with minimal performance loss and can keep agents on a malicious trajectory even after the trigger is removed.

By Wenkai Huang, Siyuan Liang, Gaolei Li, Yiming Li, Tianhao Peng, Jianhua Li, Dacheng Tao
arXiv Machine Learning
Sep 14

Robust Policy Optimization via Adversarial Importance Sampling

The paper introduces Adversarial Importance Sampling (Advis), a technique that leverages importance sampling over standard training trajectories to estimate and optimize worst‑case returns without extra environment interactions or auxiliary networks, thereby capturing long‑term robustness. It also presents advrl, a modular PyTorch library that consolidates existing robustness methods and adversarial attacks into single‑file implementations for easier prototyping and reproducible evaluation. Finally, the authors highlight that optimal adversarial hyperparameters do not transfer across agents, prompting evaluation against a broader set of attackers (6–14× more configurations) and demonstrate the effectiveness of their approach on continuous control tasks.

By Amine Andam, Jamal Bentahar, Mustapha Hedabou