OpenAI Blog

AI safety via debate

We’re proposing an AI safety technique which trains agents to debate topics with one another, using a human to judge who wins.

arXiv AI
Aug 20

A Theory of Post-hoc Debate Judgement

The paper proposes a theory for judging post-hoc debates in AI, focusing on properties like reproducibility, robustness, groundedness, and explainability. It evaluates two debate‑judgement methods—LLM judges and formal computational argumentation semantics—finding similar accuracy but noting that argumentation semantics offers stronger formal guarantees. The study suggests that argumentation semantics is a preferable framework for principled debate judges in AI systems.

By Xiang Yin, Adam Dejl, Antonio Rago, Lihu Chen, Francesca Toni
OpenAI Blog
Jun 21, 2016

Concrete AI safety problems

We (along with researchers from Berkeley and Stanford) are co-authors on today’s paper led by Google Brain researchers, Concrete Problems in AI Safety. The paper explores many research problems around ensuring that modern machine learning systems operate as intended.

OpenAI Blog
Feb 19, 2019

AI safety needs social scientists

We’ve written a paper arguing that long-term AI safety research needs social scientists to ensure AI alignment algorithms succeed when actual humans are involved. Properly aligning advanced AI systems with human values requires resolving many uncertainties related to the psychology of human rationality, emotion, and biases.

arXiv AI
Jul 23

Avoiding Obfuscation with Prover-Estimator Debate

arXiv:2506. 13609v2 Announce Type: replace Abstract: Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks.

By Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras, Lijie Chen, Jiawei Li, Zhiyang Xun
arXiv AI
Sep 15

Math for AI safety: an invitation for mathematicians

The article "Math for AI safety: an invitation for mathematicians" calls for new mathematical tools to ensure AI remains understandable, controllable, and cooperative. It outlines specific mathematical fields—logic, game theory, probability, algebra, representation theory, analysis, and geometry—each paired with an open problem tailored for mathematicians without AI safety background. The piece invites researchers to contribute to designing AI that is legible, steerable, and aligned with human values.

By Lionel Levine
arXiv AI
Sep 24

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

The paper introduces Safety Nudges, a browser-based tool that displays lightweight, in situ flags when a conversational AI exhibits risky behavior such as hallucination or overconfidence. In a two‑week field study with 45 frequent chatbot users, participants reported that the nudges were useful, clear, and minimally disruptive, and most felt more aware of potential AI harms. However, increased awareness did not automatically translate into measurable changes in user behavior, underscoring the need for relevance, calibration, and user control in nudge design.

By Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
arXiv AI
6d ago

OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas

OmouAI is an interactive deliberation system that combines large language models with computational argumentation to facilitate policy debates involving humans and simulated personas such as stakeholders, experts, or devil’s advocates. Each persona generates its own arguments, which are assembled into a shared argumentation framework that users can contest, add to, or revise, ensuring human oversight. The system evaluates arguments using deterministic argumentative semantics against external goals like the UN Sustainable Development Goals, providing faithful explanations and indicating how policy recommendations affect those goals.

By Stylianos Loukas Vasileiou, Antonio Rago, William Yeoh, Georgina Curto