arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.
By Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu
Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution. Existing defenses report near-zero attack success rate on static benchmarks, yet recent adaptive evaluations show that these results collapse once the attacker is allowed to optimize against the deployed defense.
arXiv:2606. 15441v1 Announce Type: cross Abstract: Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution.
By Lipeng He, Yihan Wang, Jiawen Zhang, N. Asokan
arXiv:2607. 02514v1 Announce Type: new Abstract: As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions.
By Josh Hills, Ida Caspary, Asa Cooper Stickland
arXiv:2603. 13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents.
By Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia
DriftNet is a dual‑head trajectory Transformer designed to detect and localize prompt injection attacks in large language model agents. It processes logged tool‑call trajectories, classifying each as compromised or not while labeling every step as benign, injection point, hijacked, or failed injection. On the AgentDrift benchmark, DriftNet achieves high accuracy, with an F1 score of 0.983, 98.7% exact injection‑point recovery, and low false‑alarm rates.
By Asif Pinjari, Mithun Paul Saint-Germain
The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.
By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy
arXiv:2606. 30602v1 Announce Type: cross Abstract: Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows.
By Kunyang Li, Kyle Domico, Jonathan Gregory, Patrick McDaniel
arXiv:2610.00400v1 Announce Type: cross
Abstract: Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or...
By Haoyu Wang, Wei Zhao, Yedi Zhang, Christopher M. Poskitt, Jun Sun
The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.
By Xiaoyan Li, Yunli Wang
arXiv:2608. 12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats.
By Jiajun Ruan, Peiyang Li, Yukun Chen, Fengting Li, Chao Feng
Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.