arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.
By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland
arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.
By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
arXiv:2601. 09753v2 Announce Type: replace-cross Abstract: AI science evaluation tools aim to assess research credibility.
By Carole J. Lee
arXiv:2511. 19735v2 Announce Type: replace-cross Abstract: Randomized controlled trials (RCTs)have been the cornerstone of clinical evidence; however, their cost, duration, and restrictive eligibility criteria limit power and external validity.
By Shu Yang, Margaret Gamalo, Haoda Fu
arXiv:2605. 08827v2 Announce Type: replace Abstract: The safety of mental health AI is often judged at the wrong temporal scale.
By Srimonti Dutta, Ratna Kandala
arXiv:2606. 02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention design.
By Junjie Luo, Ritu Agarwal, Gordon Gao
AI Morbidity and Mortality (AI M&M) is a structured, blameless framework designed to review clinical AI failures. It combines standardized case intake, evidence preservation, investigator reconstruction, tool‑in‑loop attribution, and corrective‑action tracking, classifying each event across four linked dimensions: Trigger, Mechanism, Clinical Pathway, and Corrective Action. The authors demonstrate the framework with five outpatient medication and clinical decision‑support cases, achieving full agreement among reviewers on all classification axes.
By Paulius Mui, Dean F. Sittig, Steve Labkoff, Sanjay Basu
arXiv:2608.29478v1 Announce Type: cross
Abstract: Scholarly work which aims to describe potential societal impacts (e.g., risks) of proliferating technology (especially related to artificial intellig...
By Kyra Wilson, Sabrina Kang, Saloni Dash, Aylin Caliskan
arXiv:2608. 12360v1 Announce Type: cross Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks.
By Ahmed M Salih, Oliver D\'iaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir
arXiv:2609.13642v1 Announce Type: new
Abstract: We argue that a recurring failure in the evaluation of deployed AI systems occurs when data collected for operational monitoring or regulatory complian...
By Hung-Yu Lin, Xingran Huang, Qiming Guo, Jinwen Tang
arXiv:2606. 11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments.
By Michelle Vaccaro
The paper introduces the Discovery Certification Protocol (DCP), a framework that transforms claims from AI research agents into executable tests for recovery and feedback. DCP includes multiple gates that validate improvements, provide controlled information, and measure the impact of truthful feedback, while its core requires strict controls and finite‑sample bounds. Controlled audits in SQLite optimization and virtual catalyst control demonstrated zero recoveries across 96 episodes, with rigorous verification by a deterministic, LLM‑free verifier.
By Jingjie Ning, Shanshan Zhong, Xiaochuan Li, Ji Zeng