arXiv AI

Good Benchmarks

arXiv:2607. 12217v1 Announce Type: new Abstract: Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons.

arXiv AI
Aug 25

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The paper introduces Lit2Test, a benchmark that evaluates language models’ research idea proposals by requiring each idea to include a falsifiable outcome, thereby making quality decidable. Built from 200 real-paper neighborhoods, the benchmark gathers proposals from four frontier models and compares them via 1,200 blind pairwise judgments, with reliability checks and human calibration. The results show a consistent ranking of the models, driven by test and metric quality rather than fluency, and the authors release the benchmark and related artifacts for public use.

By Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Aomufei Yuan (Peking University), Yiran Yao (Tianjin University), Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Hongyao Zuo (Tianjin University), Ziwen Gong (Hainan University), Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Tong Yang (Peking University), Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University), Xiaohui Li (Huawei Technologies), Haoli Bai (Huawei Technologies)
arXiv AI
Jul 7

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

arXiv:2501. 10711v5 Announce Type: replace-cross Abstract: Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities.

By Jialun Cao, Yuk-Kit Chan, Zixuan Ling, Wenxuan Wang, Shuqing Li, Mingwei Liu, Ruixi Qiao, Yuting Han, Chaozheng Wang, Boxi Yu, Pinjia He, Shuai Wang, Zibin Zheng, Michael R. Lyu, Shing-Chi Cheung
arXiv AI
Aug 26

BenchBench-Protocol: Evaluating Real-World Wet-Lab Protocol Reasoning and Modification

BenchBench-Protocol is a new benchmark for large language models that evaluates their ability to modify real-world wet‑lab protocols. It consists of 149 protocol‑modification tasks derived from actual changes scientists made to published protocols across 96 source protocols in nine wet‑lab biology domains. The benchmark includes weighted rubric elements for correct responses and has been reviewed by domain experts to ensure high quality.

By Aditya Sivakumar, Ashu Singhal, Nicholas Larus-Stone, Nithin Parsan
arXiv AI
Jun 26

Life After Benchmark Saturation: A Case Study of CORE-Bench

arXiv:2606. 26158v1 Announce Type: new Abstract: When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version.

By Nitya Nadgir, Sayash Kapoor, Kangheng Liu, Peter Kirgis, Matilda Orona, Stephan Rabanser, Tilman Bayer, Abhishek Shetty, Yue Ling, Derrick Chan-Sew, Rumi Nakagawa, Saiteja Utpala, Zachary S. Siegel, Arvind Narayanan
arXiv Machine Learning
Jul 31

ClawTrack: Towards Trace-Level Evaluation and Improvement of Real-World Autonomous Agents

arXiv:2607. 28037v1 Announce Type: new Abstract: As LLM-based agents are deployed in complex, multi-step workflows, a critical evaluation gap has emerged: most existing benchmarks judge only final outcomes, unable to distinguish reliable reasoning from lucky success or attribute failures to specific process deficiencies, hindering attribution in long-horizon tasks.

By Xingjian Wu, Xuhang Zhu, Xingchen Liu, Junlin Liu, Jianing Wang, Linsen Guo, Xiaoyu Li, Xuezhi Cao, Xunliang Cai