arXiv AI By Ivan Bercovich

Good Benchmarks

Read the original on arXiv AI →

arXiv:2607. 12217v1 Announce Type: new Abstract: Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 7

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

arXiv:2501. 10711v5 Announce Type: replace-cross Abstract: Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the community interprets model capabilities.

By Jialun Cao, Yuk-Kit Chan, Zixuan Ling, Wenxuan Wang, Shuqing Li, Mingwei Liu, Ruixi Qiao, Yuting Han, Chaozheng Wang, Boxi Yu, Pinjia He, Shuai Wang, Zibin Zheng, Michael R. Lyu, Shing-Chi Cheung