OpenAI Blog

Our approach to alignment research

We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.

arXiv AI
Sep 3

Automated Researchers Can Mitigate Well-characterized Alignment Failures

The paper investigates whether automated alignment researchers (AARs) can post‑train language models to reduce well‑characterized alignment failures such as deception, sycophancy, and jailbreaks while preserving general capability. Across ten failures, the strongest AAR methods significantly lower targeted failures and generalize to held‑out benchmarks, larger models, and multi‑turn audits. In contrast, a human baseline of 28 experienced researchers, given eight hours to devise one‑shot methods, underperformed the best AAR approaches, and providing human ideas to AARs did not improve results.

By Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner
arXiv AI
Jul 17

Align AI to Dynamic Human-AI Workflows

arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.

By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv Computation and Language
Aug 31

AI Alignment through a Game-theoretic Lens: A Survey

The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.

By Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang