The paper introduces a three‑layer security framework designed to protect retrieval‑augmented generation (RAG) chatbots from both direct and indirect prompt injection attacks. Layer 1 filters user input with rule‑based patterns and a semantic anomaly classifier; Layer 2 enforces a provenance‑based instruction hierarchy during context assembly; Layer 3 audits model output with a policy rule engine and semantic drift detector. Evaluations on GPT‑4o, Llama 3, and Mistral 7B demonstrate a reduction in attack success rate from 71.4 % to 11.3 %, outperforming existing single‑layer defenses while keeping false positives low and latency acceptable.
By Gulshan Saleem, Nisar Ahmed, Muhammad Imran Zaman, Ali Hassan, Umar Mujahid
arXiv:2608. 08027v1 Announce Type: cross Abstract: Prompt injection is a critical security threat in large language model (LLM) applications, where attackers hijack model behavior by embedding malicious instructions in user or external data.
By Laiqiao Qin, Tianqing Zhu, Longxiang Gao, Wanlei Zhou
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich
arXiv:2603. 19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks.
By Shawn Li, Yue Zhao
Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.
By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov
arXiv:2602. 09222v2 Announce Type: replace-cross Abstract: Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sites and performing actions on users' behalf.
By Georgios Syros, Evan Rose, Brian Grinstead, Christoph Kerschbaumer, William Robertson, Cristina Nita-Rotaru, Alina Oprea
arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.
By Hiskias Dingeto, Will Leeney
arXiv:2608. 07808v1 Announce Type: cross Abstract: Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than structured exploits, despite advancing agent capabilities and threat actors embedding injections to subvert AI-assisted security analysis.
By Jeremy McHugh
The paper introduces a prompt‑injection detection framework for email assistants that models attacks as a chain of stages. It combines a text detector, stage‑specific verifiers, rule‑based risk signals, user intent consistency checks, and a logistic decision policy. Experiments on five benchmarks show the framework outperforms pretrained detectors, achieving a mean F1 of 0.406 versus 0.216, and demonstrate that training on benign emails resembling attacks reduces false alarms.
By Ahmad Hashmi, Dhyey Patel, Yunting Yin
WAInjectBench introduces the first comprehensive benchmark for detecting prompt injection attacks against web agents, offering a fine‑grained categorization of threats and datasets that include malicious and benign text and image samples. The study systematically evaluates both text‑based and image‑based detection methods across multiple scenarios, revealing that detectors perform well on attacks with explicit instructions or visible perturbations but struggle with subtle or instruction‑free attacks. The authors release the datasets and code to facilitate further research in this area.
By Yinuo Liu, Xilong Wang, Ruohan Xu, Yuqi Jia, Neil Zhenqiang Gong
arXiv:2607. 26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools.
By Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng
The paper investigates prompt injection attacks on 14 open‑source and 3 closed‑source large language models (LLMs), introducing a new metric called Attack Success Probability (ASP) that accounts for uncertainty in model responses. It demonstrates that a simple hypnotism attack can trigger objectionable behavior in models such as StableLM2, Mistral, Openchat, and Vicuna, achieving roughly 90% ASP. The study highlights that moderately well‑known LLMs are particularly vulnerable, underscoring the importance of public awareness and effective mitigation strategies.
By Jiawen Wang, Pritha Gupta, Eyke H\"ullermeier, Xiaoxue Gao, Nancy F. Chen