arXiv AI

Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security

arXiv:2607. 18063v1 Announce Type: cross Abstract: LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation.

arXiv AI
Jun 26

Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents

arXiv:2606. 26479v1 Announce Type: cross Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions.

By Praneeth Narisetty, Shiva Nagendra Babu Kore, Uday Kumar Reddy Kattamanchi, Jayaram Kumarapu
arXiv AI
Sep 2

EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities

EvoFlint is an evolutionary atlas that maps multi‑turn vulnerabilities in large language models by treating attack discovery as a search problem rather than a generation task. It uses evolutionary quality‑diversity search to evolve phased conversation plans, employing Pareto fitness for success rate and severity, novelty search for diversity, and a generation‑level memory to incorporate model insights. The resulting risk‑indexed archive, tested on HarmBench, shows high attack success rates across models such as Claude Sonnet, GPT‑5, and Qwen3, revealing which harm categories each model’s safety training covers or misses.

By Feitong Qiao, Liren Peng, Shiming Ren, Aishwarya Jadhav, Arghavan Bahadorinejad, Marinette Chen, Muhan Zhang, Abdulaziz Suria, Gennevi Lu, Anish Das Sarma
arXiv AI
Aug 26

ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models

The paper introduces ADVERSA, an automated red‑teaming framework that evaluates large language model safety over multiple turns by tracking continuous compliance trajectories instead of binary jailbreak outcomes. Using a fine‑tuned 70B attacker model and a structured 5‑point rubric, the authors conduct controlled experiments on three frontier victim models, measuring guardrail degradation and judge reliability through a triple‑judge consensus. Results show a 26.7% jailbreak rate with most breaches occurring early, and the study documents inter‑judge agreement, attacker drift, and attacker refusals as key factors affecting safety assessment.

By Harry Owiredu-Ashley
arXiv Machine Learning
Sep 10

CoRL: Co-Evolutionary Reinforcement Learning for Adaptive Indirect Prompt-Injection Attacks and Defenses

The paper introduces CoRL, a co-evolutionary reinforcement learning framework designed to defend against adaptive indirect prompt-injection attacks on tool-augmented language agents. CoRL operates in three stages—attacker fine‑tuning, bilateral Co‑PPO training, and defender fine‑tuning—using verifier‑grounded repairs to adapt to changing attack strategies. Experiments on 1,514 executions show that CoRL reduces attack success rates to 0% while improving task utility, demonstrating its effectiveness against adaptive adversaries.

By Boyang Zhang, Qingxin Xiao, Lingwei Dang, Qingyao Wu
arXiv Machine Learning
4d ago

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

The paper introduces a curriculum reinforcement learning approach to overcome the cold‑start problem in prompt‑injection red‑teaming of frontier large language models. By training an attacker LLM sequentially against increasingly robust target models and ensuring partial success at each stage, the method achieves high attack success rates (93.8% against GPT‑5.6‑Luna and 45.0% against GPT‑5.6‑Terra) where prior RL methods fail. The attacker LLM also transfers its effectiveness to other frontier models it was not explicitly trained on.

By Chenlong Yin, Xiaolong Jin, Wei Zou, Yanting Wang, Jinyuan Jia
arXiv Computation and Language
Aug 31

Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers

The paper investigates whether stacking multiple defenses around large language models (LLMs) truly compounds security. Using the Adversary Access‑Tier Model (AATM) and a cost‑tiering system, the authors analyze a seven‑layer defense stack and find that failure correlations between layers are consistently positive, meaning the residual attack success is higher than the multiplicative prediction. Despite high coverage and low false refusals, the stack’s performance is largely driven by common architectural causes rather than diverse, independent defenses.

By Abrar Alotaibi, Muhammad Shahid Jabbar, Sadam Al-Azani, Moataz Ahmed
arXiv AI
Jun 19

LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems

arXiv:2606. 20408v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.

By Hanwool Lee, Dasol Choi, Bokyeong Kim, Seung Geun Kim, Haon Park