Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
arXiv:2607. 26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs.
By Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cer\'on Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
arXiv:2602. 17737v2 Announce Type: replace-cross Abstract: Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior.
By Upasana Biswas, Durgesh Kalwar, Subbarao Kambhampati, Sarath Sreedharan
PersonaTeaming introduces a workflow that incorporates personas into adversarial prompt generation for generative AI, achieving higher attack success rates than the state‑of‑the‑art RainbowPlus while preserving prompt diversity. The system is extended into a user‑facing playground that lets red‑teamers create their own personas and collaborate with AI to refine prompts, fostering diverse strategies. A user study with 11 industry practitioners found the playground produced useful outputs and encouraged out‑of‑the‑box thinking, even when suggestions were not strictly followed.
By Wesley Hanwen Deng, Mingxi Yan, Sunnie S. Y. Kim, Akshita Jha, Lauren Wilcox, Kenneth Holstein, Motahhare Eslami, Leon A. Gatys