The paper discusses the need to adapt incident reporting frameworks for AI agents, which are rapidly deployed and face unique security challenges. By comparing AI systems and agents and consulting 23 experts, the authors identify key reporting elements such as agent memory, autonomy levels, and tool usage. They also highlight open research questions, potential reporting weaknesses like data leakage, and outline privacy requirements for secure AI agent deployment.
By Anastasia Pustozerova, Eugene Bagdasarian, Luca Beurer-Kellner, Battista Biggio, Nico Ebert, David Filip, Marc Fischer, Heather Frase, David Hofer, Juliane Hoffmann, Daphne Ippolito, Somesh Jha, Sean McGregor, Esfandiar Mohammadi, Luca Nannini, Cristina Nita-Rotaru, Alina Oprea, Kevin Paeth, Andrew Paverd, Jonathan Petit, Andreas Rauber, Christian Riess, John Sotiropoulos, Andreas Wespi, Kathrin Grosse
arXiv:2608. 04921v1 Announce Type: cross Abstract: As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address risks arising from interactions among system components and deployment environments.
By Leah Davis, Dominic Martin, AJung Moon
The paper introduces the AI Assessment Sandbox Configurator, an open‑source framework designed to support technical assessment in AI Regulatory Sandboxes (AIRS) mandated by the EU Artificial Intelligence Act. It outlines 11 architectural and governance requirements for infrastructure that enables large‑scale, structured technical testing, and presents a catalogue of tests, a shared data model, dashboards, and reporting tools that harmonise heterogeneous outputs. An early‑stage pilot demonstrated the framework’s harmonisation and reporting capabilities within a live AIRS engagement, contributing to an official Exit Report.
By Alessio Buscemi, German Castignani, Daniele Pagani, Maxime Cordy, Jordi Cabot
arXiv:2603. 26270v2 Announce Type: replace-cross Abstract: Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging because many vulnerabilities are tightly coupled with project-specific business logic.
By Ziqiao Kong, Wanxu Xia, Chong Wang, Yue Xue, Yi Lu, Pan Li, Shaohua Li, Zong Cao, Yang Liu
arXiv:2607. 05163v1 Announce Type: cross Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate.
By Harleen Kaur Sidhu, Rebecca Scholefield, Nour Annan, Kevin Hernandez, Isabel Nieh Hou, Abdulrahman Alshaikhi, Ze Shen Chin, Rokas Gipi\v{s}kis
arXiv:2608. 07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks.
By Afreen Alam, Evgenija Popchanovska, Ana Gjorgjevikj, Maryan Rizinski, Lubomir T. Chitkushev, Irena Vodenska, Dimitar Trajanov