arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.
By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland
arXiv:2603. 11001v3 Announce Type: replace-cross Abstract: Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or similar methodologies, increasingly inform frontier AI governance and deployment decisions.
By Patricia Paskov, Kevin Wei, Shen Zhou Hong, Dan Bateyko, Xavier Roberts-Gaal, Carson Ezell, Gailius Praninskas, Valerie Chen, Umang Bhatt, Ella Guest
arXiv:2601. 09753v2 Announce Type: replace-cross Abstract: AI science evaluation tools aim to assess research credibility.
By Carole J. Lee
arXiv:2511. 19735v2 Announce Type: replace-cross Abstract: Randomized controlled trials (RCTs)have been the cornerstone of clinical evidence; however, their cost, duration, and restrictive eligibility criteria limit power and external validity.
By Shu Yang, Margaret Gamalo, Haoda Fu
arXiv:2605. 08827v2 Announce Type: replace Abstract: The safety of mental health AI is often judged at the wrong temporal scale.
By Srimonti Dutta, Ratna Kandala
arXiv:2606. 02458v1 Announce Type: new Abstract: Organizations routinely run experiments for A/B testing, yet the data generated from one experiment is underutilized to inform subsequent intervention design.
By Junjie Luo, Ritu Agarwal, Gordon Gao
arXiv:2608. 12360v1 Announce Type: cross Abstract: Background: AI/ML-enabled medical devices are increasingly deployed in healthcare under evolving regulatory frameworks.
By Ahmed M Salih, Oliver D\'iaz, Alejandro Guzman, Noah Marquez Vara, Fotios Avgoustidis, Rituraj Singh, Saman Barakat, Zahra Raisi-Estabragh, Karim Lekadir
arXiv:2606. 11217v1 Announce Type: cross Abstract: The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: "in silico" behavioral experiments.
By Michelle Vaccaro
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
arXiv:2608. 14598v1 Announce Type: new Abstract: Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions.
By Shiva Kaul, Anjum Khurshid
arXiv:2608. 15417v1 Announce Type: cross Abstract: Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used.
By Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory, Sreeparvathy Sajeev, Utkarsh Tomar, Avyay M Casheekar
arXiv:2608. 07202v1 Announce Type: new Abstract: Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines.
By Milan Markovic, Goutham Indukuri, Somayajulu Sripada, Colby J. Vorland, Jack Wilkinson, Clare Robertson, Mark Bolland, Andrew Grey, Miriam Brazzelli, Alison Avenell