Hugging Face Trending Papers
Jun 13

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.

arXiv Machine Learning
Sep 16

CoER: Defending against Adaptive Indirect Prompt Injection via Adversarial Co-Evolution and Refinement

The paper introduces CoER, a framework that defends language‑model agents against adaptive indirect prompt injection (IPI) by employing attacker‑defender co‑evolution and refinement. CoER models IPI as a general‑sum Markov game, uses Co‑PPO to maintain historical opponent populations, and fine‑tunes defenders only on verified safe demonstrations. In experiments across seven domains, CoER cuts attack success from 38.5% to 0.2% while boosting task utility from 63.2% to 76.3%.

By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
arXiv AI
Sep 30

Render Before Reading: Visual Rendering as a Prompt Injection Defense

The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.

By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r