arXiv Machine Learning

BadWAM: When World-Action Models Dream Right but Act Wrong

arXiv:2607. 15207v1 Announce Type: new Abstract: World-action models (WAMs) are emerging as a promising foundation for embodied control: rather than predicting actions alone, they learn representations that couple action generation with future world prediction.

arXiv Machine Learning
Sep 10

TrojanWorld: Backdooring World-Model Agents via Imagination Steering

TrojanWorld is a backdoor framework that targets world-model agents by steering their internal imagination toward attacker-specified actions when a physical trigger is present. The attack uses Decision-Reflective Induction, Clean Behavior Anchoring, and Causal Propagation to maintain stealth, persistence, and high performance. Experiments on TD-MPC2, DreamerV3, and R2-Dreamer across several benchmarks show that the attack can induce target actions with minimal performance loss and can keep agents on a malicious trajectory even after the trigger is removed.

By Wenkai Huang, Siyuan Liang, Gaolei Li, Yiming Li, Tianhao Peng, Jianhua Li, Dacheng Tao
arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai
Hugging Face Trending Papers
Jun 13

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.