arXiv AI

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

arXiv:2605. 02050v2 Announce Type: replace-cross Abstract: This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies).

arXiv AI
Jun 19

Measuring Biological Capabilities and Risks of AI Agents

arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.

By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland
arXiv AI
Jul 1

RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies

arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.

By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
arXiv AI
Sep 2

AI Morbidity and Mortality: A Framework for Clinical AI Failure Review

AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.

By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv AI
Sep 11

Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents

The paper introduces the Discovery Certification Protocol (DCP), a framework that transforms claims from AI research agents into executable tests for recovery and feedback. DCP includes multiple gates that validate improvements, provide controlled information, and measure the impact of truthful feedback, while its core requires strict controls and finite‑sample bounds. Controlled audits in SQLite optimization and virtual catalyst control demonstrated zero recoveries across 96 episodes, with rigorous verification by a deterministic, LLM‑free verifier.

By Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng