Advancing red teaming with people and AI
Read the original on OpenAI Blog →Advancing red teaming with people and AI
Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.
Advancing red teaming with people and AI
Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
arXiv:2607. 26115v1 Announce Type: cross Abstract: We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs.
arXiv:2602. 17737v2 Announce Type: replace-cross Abstract: Mutual adaptation is a central challenge in human-AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior.
arXiv:2607. 02198v1 Announce Type: cross Abstract: Human-AI teaming has received increasing attention in the literature.
arXiv:2606. 10906v1 Announce Type: cross Abstract: We study models for human-AI teaming through the lens of statistical calibration.
arXiv:2608. 13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.