arXiv:2606. 15441v1 Announce Type: cross Abstract: Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution.
By Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan
Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.
The paper introduces CoER, a framework that defends language‑model agents against adaptive indirect prompt injection (IPI) by employing attacker‑defender co‑evolution and refinement. CoER models IPI as a general‑sum Markov game, uses Co‑PPO to maintain historical opponent populations, and fine‑tunes defenders only on verified safe demonstrations. In experiments across seven domains, CoER cuts attack success from 38.5% to 0.2% while boosting task utility from 63.2% to 76.3%.
By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
arXiv:2608. 09542v1 Announce Type: cross Abstract: Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs.
By Hongli Shen, Shaopeng Fu, Qinbo Zhang, Jian Li, Di Wang
The paper investigates how multimodal large language models are more susceptible to prompt injection when adversarial instructions are presented as text rather than as non-textual inputs like images. It proposes a training‑free defense that renders untrusted payloads into typographic images (or audio) before they reach the model, a method called Pictionary. Experiments on ten models and two benchmarks show that this approach significantly lowers attack success rates while maintaining normal functionality, and that fine‑tuning on image‑rendered instructions can further reduce the modality gap.
By Jie Zhang, Andrei Baroian, Jan N. van Rijn, Avital Shafran, Florian Tram\`{e}r
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e. g.
arXiv:2609.22234v1 Announce Type: cross
Abstract: Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in...
By Nicholas Sansoterra, Zishuo Zheng, Sachin Kumar
DiSCO is a zero‑shot, black‑box defense for text‑to‑image models that operates solely at the prompt level. It expands prompts with a distribution‑guided suffix using beam search and contrastive scoring against safe and unsafe image pools generated by the target model, iteratively refining until safe content is produced. The method improves safety on the I2P benchmark under various red‑teaming attacks, reducing attack success rates by 37.7% and 25.13% while preserving semantic fidelity and image coherence.
By Tong Zhang, Motasem Alfarra, Carlos Hinojosa, Christos Louizos, Bernard Ghanem
arXiv:2512. 20806v3 Announce Type: replace Abstract: Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment.
By Anselm Paulus, Ilia Kulikov, Brandon Amos, R\'emi Munos, Ivan Evtimov, Kamalika Chaudhuri, Arman Zharmagambetov
The paper introduces MINT‑Safe, a new open‑source dataset of 11,270 multi‑image dialogues and 500 refusal VQA pairs designed to expose safety risks in multi‑modal large language models during open‑ended conversations. It also proposes TAD‑Align, a turn‑aware dual‑objective reward framework that dynamically up‑weights dialogue turns with inconsistent safety behavior, improving safety metrics on Qwen2.5‑VL‑7B‑Instruct and LLaVA‑Next‑7B. The results show over 10% reduction in attack success rate and notable gains in harmlessness and helpfulness while maintaining overall model performance.
By Han Zhu, Jiale Chen, Chengkun Cai, Shengjie Sun, Haoran Li, Yujin Zhou, Chi-Min Chan, Pengcheng Wen, Lei Li, Yike Guo, Sirui Han
arXiv:2604. 00310v2 Announce Type: replace-cross Abstract: Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions.
By Anurag Kumar, Raghuveer Peri, Jon Burnsky, Alexandru Nelus, Rohit Paturi, Srikanth Vishnubhotla, Yanjun Qi
arXiv:2603. 19423v2 Announce Type: replace-cross Abstract: Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks.
By Shawn Li, Yue Zhao