OpenAI Blog

Disrupting malicious uses of AI: June 2025

Our latest report featuring case studies of how we’re detecting and preventing malicious uses of AI.

OpenAI Blog
Feb 20, 2018

Preparing for malicious uses of AI

We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.

arXiv AI
Sep 4

Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression

The article proposes a structured framework of behavioral indicators that could signal a progression toward potentially catastrophic threats from AI systems. Drawing on established methods from cybersecurity and national security, it defines clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The framework is intended to enable researchers and policymakers to implement evidence‑based monitoring protocols for rogue AI progression.

By T. Bauer, W. P. Kegelmeyer, E. Begoli, A. Sadovnik, T. Emerson, C. Corley, N. Generous, J. Moore, B. Bartoldson, R. Goldhan, M. Goldman, M. Greaves, M. J. D. Vermeer, B. MacLennan, D. Schulker, N. VanHoudnos, J. Bansemer, Y. Bengio
arXiv AI
3d ago

AI Security Research Should Better Incentivize Defense Research

The article discusses a notable imbalance in AI security research, where studies on attacking AI systems outnumber those on defending them. It highlights that this skew is evident across various subfields such as federated learning, speech recognition, membership inference, and large language models. The authors argue that attack papers often benefit from favorable evaluation conditions, whereas defense papers face stricter standards, resulting in a literature rich in vulnerabilities but lacking robust, deployable protections.

By Youqian Zhang
Hugging Face Trending Papers
Jul 8

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks. These technologies power real-time threat detection, phishing defense, secure code generation, and vulnerability exploitation at unprecedented scales.

arXiv AI
Jul 9

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

arXiv:2607. 06963v1 Announce Type: cross Abstract: Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks.

By Kiarash Ahi, Saeed Valizadeh
arXiv AI
Sep 23

Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

The paper discusses the need to adapt incident reporting frameworks for AI agents, which are rapidly deployed and face unique security challenges. By comparing AI systems and agents and consulting 23 experts, the authors identify key reporting elements such as agent memory, autonomy levels, and tool usage. They also highlight open research questions, potential reporting weaknesses like data leakage, and outline privacy requirements for secure AI agent deployment.

By Anastasia Pustozerova, Eugene Bagdasarian, Luca Beurer-Kellner, Battista Biggio, Nico Ebert, David Filip, Marc Fischer, Heather Frase, David Hofer, Juliane Hoffmann, Daphne Ippolito, Somesh Jha, Sean McGregor, Esfandiar Mohammadi, Luca Nannini, Cristina Nita-Rotaru, Alina Oprea, Kevin Paeth, Andrew Paverd, Jonathan Petit, Andreas Rauber, Christian Riess, John Sotiropoulos, Andreas Wespi, Kathrin Grosse