arXiv AI
Jun 19

Measuring Biological Capabilities and Risks of AI Agents

arXiv:2606. 19899v1 Announce Type: cross Abstract: This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities and risks of AI scientists, or agentic AI systems capable of autonomously or collaboratively performing multi-step scientific tasks.

By Patricia Paskov, Jeffrey Lee, Kyle Brady, Alyssa Worland
arXiv AI
Jul 22

BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment

arXiv:2607. 05462v2 Announce Type: replace-cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse.

By Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
arXiv AI
Jul 8

Evaluating calibrated refusal and safe usefulness in dual-use biology settings

arXiv:2607. 05462v1 Announce Type: cross Abstract: As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse.

By Edwin H. Wintermute, Harmon Bhasin, Christina M. Agapakis, Dianzhuo Wang, Evan Seeyave, Arjun Banerjee, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle, Kenny Workman
arXiv AI
Aug 19

Traceable Trust for action-ready artificial intelligence in bioscience

The article introduces Traceable Trust, a framework designed to guide the transition from AI-generated outputs to laboratory actions in bioscience. It outlines a reviewable process that evaluates evidence, claimed capabilities, delegated agency, action thresholds, override mechanisms, and feedback loops. Three case studies demonstrate how the framework can document trust as AI outputs influence scientific work.

By Huayu Xin, Yizhi Cai, Mukilan Deivarajan Suresh, Gavin Michael Farrell, Iwona Gajda, Charlie Harrison, Conor Houghton, Mato Lagator, Yang Lu, Virginia Portillo, Reyer Zwiggelaar, Sebastian Lobentanzer