arXiv AI By Charles Pozniak, Jeba Sania

The Foreign Policy AI Evaluation Gap

Read the original on arXiv AI →

arXiv:2607. 02955v1 Announce Type: cross Abstract: We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

The AI Assessment Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

The paper introduces the AI Assessment Sandbox Configurator, an open‑source framework designed to support technical assessment in AI Regulatory Sandboxes (AIRS) mandated by the EU Artificial Intelligence Act. It outlines 11 architectural and governance requirements for infrastructure that enables large‑scale, structured technical testing, and presents a catalogue of tests, a shared data model, dashboards, and reporting tools that harmonise heterogeneous outputs. An early‑stage pilot demonstrated the framework’s harmonisation and reporting capabilities within a live AIRS engagement, contributing to an official Exit Report.

By Alessio Buscemi, German Castignani, Daniele Pagani, Maxime Cordy, Jordi Cabot
arXiv AI
Jun 19

Measuring Biological Capabilities and Risks of AI Agents

arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.

By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland