arXiv AI

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

arXiv:2607. 06608v1 Announce Type: cross Abstract: We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.

arXiv AI
Sep 23

Beyond Predictable Paths: Redefining AI Security Incident Reporting for Agents

The paper discusses the need to adapt incident reporting frameworks for AI agents, which are rapidly deployed and face unique security challenges. By comparing AI systems and agents and consulting 23 experts, the authors identify key reporting elements such as agent memory, autonomy levels, and tool usage. They also highlight open research questions, potential reporting weaknesses like data leakage, and outline privacy requirements for secure AI agent deployment.

By Anastasia Pustozerova, Eugene Bagdasarian, Luca Beurer-Kellner, Battista Biggio, Nico Ebert, David Filip, Marc Fischer, Heather Frase, David Hofer, Juliane Hoffmann, Daphne Ippolito, Somesh Jha, Sean McGregor, Esfandiar Mohammadi, Luca Nannini, Cristina Nita-Rotaru, Alina Oprea, Kevin Paeth, Andrew Paverd, Jonathan Petit, Andreas Rauber, Christian Riess, John Sotiropoulos, Andreas Wespi, Kathrin Grosse
arXiv AI
Jul 31

AI Security Priorities: A Field-Wide Agenda

arXiv:2607. 26069v1 Announce Type: cross Abstract: As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen.

By Gil Gekker, Rachel Steratore, Everett Smith, Asher Brass-Gershovich, Varun Gandhi, Nicole Nichols, Vijay Bolina, Buck Shlegeris, Lisa Einstein, Dan Lahav, Omer Nevo, Sella Nevo
arXiv AI
Aug 20

Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions

The article reviews the emergence of Agentic AI, covering its evolution, theoretical foundations, working principles, and architectural aspects. It surveys recent scholarly contributions across various domains, highlighting real‑world applications, current research findings, and existing challenges. The review also proposes a framework for stakeholder adoption and outlines future research directions to guide researchers and practitioners.

By AKM Bahalul Haque, Al Amin Islam Ridoy, Mohammad Rayhan, Ivan Porres
OpenAI Blog
Feb 20, 2018

Preparing for malicious uses of AI

We’ve co-authored a paper that forecasts how malicious actors could misuse AI technology, and potential ways we can prevent and mitigate these threats. This paper is the outcome of almost a year of sustained work with our colleagues at the Future of Humanity Institute, the Centre for the Study of Existential Risk, the Center for a New American Security, the Electronic Frontier Foundation, and others.

OpenAI Blog
Apr 16, 2020

Improving verifiability in AI development

We’ve contributed to a multi-stakeholder report by 58 co-authors at 30 organizations, including the Centre for the Future of Intelligence, Mila, Schwartz Reisman Institute for Technology and Society, Center for Advanced Study in the Behavioral Sciences, and Center for Security and Emerging Technologies. This report describes 10 mechanisms to improve the verifiability of claims made about AI systems.

arXiv AI
3d ago

When Does Randomized Oversight Align AI Agents That Can Conceal?

The paper investigates how randomized audits and scoring can align AI agents that are capable of concealing misconduct and manipulating records. It finds that stronger auditing can actually make violations harder to detect, and that effective deterrence requires either sanctions beyond simple forfeiture or conditions where evidence survives concealment and audit timing is unpredictable. The study also highlights that when evidence can be erased, deterrence must rely on reducing the gains from violating or increasing the cost of concealment, and it uses the July 2026 incident involving OpenAI’s cybersecurity evaluations and Hugging Face’s infrastructure as a case study.

By Joshua S. Gans, Richard Holden