Disrupting malicious uses of AI | February 2026
Our latest threat report examines how malicious actors combine AI models with websites and social platforms—and what it means for detection and defense.
We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.
Our latest threat report examines how malicious actors combine AI models with websites and social platforms—and what it means for detection and defense.
Our latest report featuring case studies of how we’re detecting and preventing malicious uses of AI.
arXiv:2609.14796v1 Announce Type: new Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion a...
arXiv:2607. 02197v1 Announce Type: cross Abstract: The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliable AI systems.
arXiv:2608. 05173v1 Announce Type: cross Abstract: As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole.
The article proposes a structured framework of behavioral indicators that could signal a progression toward potentially catastrophic threats from AI systems. Drawing on established methods from cybersecurity and national security, it defines clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The framework is intended to enable researchers and policymakers to implement evidence‑based monitoring protocols for rogue AI progression.
OpenAI is investing in stronger safeguards and defensive capabilities as AI models become more powerful in cybersecurity. We explain how we assess risk, limit misuse, and work with the security community to strengthen cyber resilience.
Google DeepMind researches AI's harmful manipulation risks across areas like finance and health, leading to new safety measures.
arXiv:2607. 26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen.