Hugging Face Trending Papers

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Read the original on Hugging Face Trending Papers →

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 16

CoER: Defending against Adaptive Indirect Prompt Injection via Adversarial Co-Evolution and Refinement

The paper introduces CoER, a framework that defends language‑model agents against adaptive indirect prompt injection (IPI) by employing attacker‑defender co‑evolution and refinement. CoER models IPI as a general‑sum Markov game, uses Co‑PPO to maintain historical opponent populations, and fine‑tunes defenders only on verified safe demonstrations. In experiments across seven domains, CoER cuts attack success from 38.5% to 0.2% while boosting task utility from 63.2% to 76.3%.

By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
arXiv AI
Sep 16

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.

By Xiaoyan Li, Yunli Wang