arXiv:2609.22510v1 Announce Type: cross
Abstract: As LLM applications integrate with external tools, they are increasingly exposed to indirect prompt injection (IPI), where adversarial instructions a...
By Justin Szczepaniak, Elad Feldman, Naum Viner, Ben Nassi
arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.
By Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu
arXiv:2608.21500v1 Announce Type: cross
Abstract: Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inje...
By Yibo Peng, Long Lian, David Wagner, Sizhe Chen
arXiv:2603. 13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents.
By Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich
The paper introduces CoRL, a co-evolutionary reinforcement learning framework designed to defend against adaptive indirect prompt-injection attacks on tool-augmented language agents. CoRL operates in three stages—attacker fine‑tuning, bilateral Co‑PPO training, and defender fine‑tuning—using verifier‑grounded repairs to adapt to changing attack strategies. Experiments on 1,514 executions show that CoRL reduces attack success rates to 0% while improving task utility, demonstrating its effectiveness against adaptive adversaries.
By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
arXiv:2606. 15441v1 Announce Type: cross Abstract: Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution.
By Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan
arXiv:2609.22818v1 Announce Type: cross
Abstract: Memory-poisoning defenses for LLM agents are typically evaluated by their ability to prevent attacks. However, the traffic they process is rarely adv...
By Pritom Bhowmik
arXiv:2607. 18063v1 Announce Type: cross Abstract: LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation.
By Devina Jain, David Hartmann, Chuan Li
arXiv:2606. 14517v1 Announce Type: cross Abstract: LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents.
By Yuguang Zhou, Xunguang Wang, Pingchuan Ma, Zhantong Xue, Zhaoyu Wang, Shuai Wang
arXiv:2606. 06529v1 Announce Type: new Abstract: An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately.
By Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad IV, Joachim Schaeffer, Ram Potham, Tyler Tracy
arXiv:2609.38291v1 Announce Type: cross
Abstract: Large language model (LLM) agents are vulnerable to safety risks such as injected malicious instructions or misleading information, motivating runtim...
By Zhuo Liu, Moxin Li, Zhixin Ma, Wentao Shi, Wenjie Wang, Fuli Feng