When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
TrojanWorld is a backdoor framework that targets world-model agents by steering their internal imagination toward attacker-specified actions when a physical trigger is present. The attack uses Decision-Reflective Induction, Clean Behavior Anchoring, and Causal Propagation to maintain stealth, persistence, and high performance. Experiments on TD-MPC2, DreamerV3, and R2-Dreamer across several benchmarks show that the attack can induce target actions with minimal performance loss and can keep agents on a malicious trajectory even after the trigger is removed.
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
arXiv:2606. 18697v1 Announce Type: new Abstract: Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments.
arXiv:2608. 11295v1 Announce Type: cross Abstract: Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing.
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2510.17021v2 Announce Type: replace-cross Abstract: Large language model (LLM) unlearning is a key approach for removing undesired data, knowledge, or behaviors from pretrained models while ret...