OpenAI Blog

Advancing red teaming with people and AI

Read the original on OpenAI Blog →

Advancing red teaming with people and AI

Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.

arXiv Machine Learning
Jul 30

GPT-Red: Automated Red Teaming via Self-Play at Scale

arXiv:2607. 26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs.

By Eric Wallace, Christopher A. Choquette-Choo, Nikhil Kandpal, Sam Toyer, Dylan Hunn, Stephanie Lin, Yuxin Wen, Xiangyu Qi, Christopher Wolff, Zizhao Wang, Milad Nasr, Sicheng Zhu, Chuan Guo, Juan Felipe Cer\'on Uribe, Kaiwen Wang, Aiden Low, Kai Xiao, Kai Chen
arXiv AI
2d ago

AI Evaluation Should Work With Humans

arXiv:2608. 13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.

By Jan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas