arXiv AI By Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, Nithin Parsan

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

Read the original on arXiv AI →

BenchBench-Protocol is a new benchmark for large language models that evaluates their ability to modify real-world wet‑lab protocols. It consists of 149 protocol‑modification tasks derived from actual changes scientists made to published protocols across 96 source protocols in nine wet‑lab biology domains. The benchmark includes weighted rubric elements for correct responses and has been reviewed by domain experts to ensure high quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research

arXiv:2606. 07591v1 Announce Type: cross Abstract: AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify.

By Wanghan Xu, Shuo Li, Tianlin Ye, Qinglong Cao, Yixin Chen, Hengjian Gao, Yiheng Wang, Qi Li, Kun Li, Sheng Xu, Shengdu Chai, Fangchen Yu, Xiangyu Zhao, Zhangrui Zhao, Weijie Ma, Zijie Guo, Haoyu Zhou, Haoxiang Yin, Lixue Cheng, Chaofan Hu, Haoxuan Li, Lu Mi, Xuxuan Xie, Yifan Zhou, Ruizhe Chen, Zhiwang Zhou, Xingjian Guo, Yuhao Zhou, Xuming He, Shengyuan Xu, Xinyu Gu, Jiamin Wu, Mianxin Liu, Chunfeng Song, Fenghua Ling, Dongzhan Zhou, Shixiang Tang, Yuqiang Li, Mao Su, Peng Ye, Siqi Sun, Bin Wang, Xue Yang, Zhenfei Yin, Tianfan Fu, Guangtao Zhai, Wanli Ouyang, Bo Zhang, Lei Bai, Wenlong Zhang