arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv:2607. 24893v1 Announce Type: cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in its observations, spreads them across several agents, and an external step reassembles and executes them after the run.
By Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2608.30362v1 Announce Type: new
Abstract: As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Succe...
By Yunseok Lee, Yunji Kim, Woojin Lee
The paper introduces covert indirect prompt injection (IPI) attacks on tool‑using large language model agents, distinguishing between covert and overt successes. It defines new metrics—Covert Success Rate (CSR) and Overt Success Rate (OSR)—to capture whether users notice the injection. The authors propose ICoA, an attack that steers agents back to the user’s task after executing the injection, achieving higher CSR than existing methods on four target models.
By Yunseok Lee, Yunji Kim, Woojin Lee
The paper introduces skill cascading attacks, where a malicious goal is spread across multiple seemingly benign skills, causing harmful outcomes when combined. It presents SkillCascade, an automated red‑teaming framework, and releases SkillCascade‑Bench, a benchmark of 213 validated cascading test cases across various agent systems and domains. Experiments show that these cascaded interactions reliably induce harmful behaviors while evading existing per‑skill scanners and runtime monitors, revealing a gap between component‑level integrity and system‑level safety.
By Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
arXiv:2605. 08442v5 Announce Type: replace-cross Abstract: We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not necessarily block execution, and vice versa.
By Jun Wen Leong
arXiv:2609.15781v1 Announce Type: cross
Abstract: Pretrained world models, learned simulators that encode an observation into a latent state and predict how it evolves under actions, are beginning to...
By Roberto Ria\~no, Gorka Abad, Stjepan Picek, Aitor Urbieta
arXiv:2608. 09577v1 Announce Type: new Abstract: Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisoned skill can persistently compromise every agent that installs it.
By Hao Sui, Simeng Qin, Jie Liao, Xiaojun Jia, Bing Chen, Yang Liu
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich
The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.
By Xiaoyan Li, Yunli Wang
arXiv:2609.22510v1 Announce Type: cross
Abstract: As LLM applications integrate with external tools, they are increasingly exposed to indirect prompt injection (IPI), where adversarial instructions a...
By Justin Szczepaniak, Elad Feldman, Naum Viner, Ben Nassi