The article proposes a structured framework of behavioral indicators that could signal a progression toward potentially catastrophic threats from AI systems. Drawing on established methods from cybersecurity and national security, it defines clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The framework is intended to enable researchers and policymakers to implement evidence‑based monitoring protocols for rogue AI progression.
By T. Bauer, W. P. Kegelmeyer, E. Begoli, A. Sadovnik, T. Emerson, C. Corley, N. Generous, J. Moore, B. Bartoldson, R. Goldhan, M. Goldman, M. Greaves, M. J. D. Vermeer, B. MacLennan, D. Schulker, N. VanHoudnos, J. Bansemer, Y. Bengio
arXiv:2606. 28929v1 Announce Type: cross Abstract: Cybersecurity is a real-life test-bed for many machine learning problems at once, especially when considering modern strides in using Large Language Models (LLMs) to automate processes as ``agents.
By Edward Raff, Maor Ashkenazi, Sagar Samtani, David J. Elkind, Sven Krasser
arXiv:2502. 16184v3 Announce Type: replace Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems.
By Henrik Nolte, Miriam Rateike, Mich\`ele Finck
arXiv:2606. 02644v1 Announce Type: cross Abstract: Agentic scaffolds have dramatically improved LLM performance on complex, long-horizon tasks, yielding both broad benefits and amplified risks in domains like cybersecurity.
By Eliot Krzysztof Jones, Mateusz Dziemian, Matt Fredrikson, J Zico Kolter
arXiv:2609.14796v1 Announce Type: new
Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion a...
By Joshua Levy, Mick Yang, Kellin Pelrine
Our latest report featuring case studies of how we’re detecting and preventing malicious uses of AI.