OpenAI Blog

Preparing for malicious uses of AI

We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.

arXiv AI
Sep 4

Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression

The article proposes a structured framework of behavioral indicators that could signal a progression toward potentially catastrophic threats from AI systems. Drawing on established methods from cybersecurity and national security, it defines clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The framework is intended to enable researchers and policymakers to implement evidence‑based monitoring protocols for rogue AI progression.

By T. Bauer, W. P. Kegelmeyer, E. Begoli, A. Sadovnik, T. Emerson, C. Corley, N. Generous, J. Moore, B. Bartoldson, R. Goldhan, M. Goldman, M. Greaves, M. J. D. Vermeer, B. MacLennan, D. Schulker, N. VanHoudnos, J. Bansemer, Y. Bengio
arXiv AI
Jul 31

AI Security Priorities: A Field-Wide Agenda

arXiv:2607. 26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen.

By Gil Gekker, Rachel Steratore, Everett Smith, Asher Brass-Gershovich, Varun Gandhi, Nicole Nichols, Vijay Bolina, Buck Shlegeris, Lisa Einstein, Dan Lahav, Omer Nevo, Sella Nevo