arXiv Machine Learning By Shuhao Zhang, Jiarui Li, Qi Cao, Ruiyi Zhang, Pengtao Xie

Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense

Read the original on arXiv Machine Learning →

arXiv:2605. 30837v2 Announce Type: replace-cross Abstract: Prompt-injection detectors are heterogeneous: each is strong on a different slice of attacks, and none is always reliable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 26

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.

By Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu
arXiv Machine Learning
Sep 10

CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

The paper introduces CoRL, a co-evolutionary reinforcement learning framework designed to defend against adaptive indirect prompt-injection attacks on tool-augmented language agents. CoRL operates in three stages—attacker fine‑tuning, bilateral Co‑PPO training, and defender fine‑tuning—using verifier‑grounded repairs to adapt to changing attack strategies. Experiments on 1,514 executions show that CoRL reduces attack success rates to 0% while improving task utility, demonstrating its effectiveness against adaptive adversaries.

By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu