arXiv:2405.18540v3 Announce Type: replace-cross
Abstract: Red-teaming, or identifying prompts that elicit harmful responses, is a critical step in ensuring the safe and responsible deployment of larg...
By Seanie Lee, Minsu Kim, Lynn Cherif, David Dobre, Juho Lee, Sung Ju Hwang, Kenji Kawaguchi, Gauthier Gidel, Yoshua Bengio, Esmeralda S. Whitammer, Moksh Jain
arXiv:2607. 13394v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) offer a promising alternative to reward-maximizing reinforcement learning (RL) for large reasoning models, encouraging diverse reasoning paths by matching reward distributions rather than collapsing to dominant modes.
By Xiaodong Liu, Michael Xu, Jack W. Stokes, Paul Smolensky, Doug Burger, Jianfeng Gao
arXiv:2607. 26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs.
By Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cer\'on Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
arXiv:2608. 10171v1 Announce Type: new Abstract: The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespread adoption.
By Berkay Ozcam, Irem Onen, Mehmet Fatih Amasyali, Emin Islam Tatli
arXiv:2606. 09701v1 Announce Type: cross Abstract: AI red teaming must continually adapt to evolving attackers and defenders.
By Blake Bullwinkel, Eugenia Kim, Amanda Minnich, Mark Russinovich
arXiv:2506. 07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence.
By Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Zhi Yu, Yuanpeng Tan, Chao Qian
arXiv:2607. 06109v1 Announce Type: cross Abstract: Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple $\ell_p$ perturbations but suffers from robustness trade-offs between different threats.
By Woo Jae Kim, Kyle Min, Suhyeon Ha, Joonsung Jeon, Sung-eui Yoon
arXiv:2605. 01729v2 Announce Type: replace Abstract: Generative Flow Networks (GFlowNets) learn to sample states proportional to an unnormalized reward.
By Zengxiang Lei, Ananth Shreekumar, Jonathan Rosenthal, Ruoyu Song, Alvaro A. Cardenas, Daniel J. Fremont, Dongyan Xu, Satish Ukkusuri, Z. Berkay Celik
arXiv:2606. 00801v1 Announce Type: cross Abstract: Current approaches to LLM adversarial testing suffer from coverage gaps: manual red-teaming does not scale, LLM-as-attacker methods exhibit mode collapse, and gradient-based approaches produce uninterpretable gibberish.
By Subhadip Mitra
arXiv:2606. 24589v1 Announce Type: new Abstract: Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm that resulting failures are real.
By Khanak Khandelwal (Indian Institute of Technology Jodhpur)
The paper introduces T-MAP, a trajectory‑aware evolutionary search technique designed to red‑team large language model agents by exploiting vulnerabilities that arise during multi‑step tool execution. Unlike traditional methods that focus on harmful text, T‑MAP uses execution trajectories to generate adversarial prompts that bypass safety guardrails and achieve harmful objectives through actual tool interactions. Experiments across various Model Context Protocol environments show that T‑MAP outperforms baseline methods in attack realization rate and remains effective against advanced models such as GPT‑5.2, Gemini‑3‑Pro, Qwen3.5, and GLM‑5.
By Hyomin Lee, Sangwoo Park, Yumin Choi, Sohyun An, Seanie Lee, Sung Ju Hwang
arXiv:2603. 13026v2 Announce Type: replace Abstract: Prompt injection poses serious security risks to real-world LLM applications, particularly autonomous agents.
By Chenlong Yin, Runpeng Geng, Yanting Wang, Jinyuan Jia