arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.
By Yangzhen Wu, Aaron J. Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn Song
arXiv:2607. 02854v1 Announce Type: cross Abstract: Before fixing an issue, it is useful to first reproduce it by generating a bug reproduction test (BRT).
By Toufique Ahmed, Jatin Ganhotra, Avraham Shinnar, Martin Hirzel
arXiv:2605. 28556v2 Announce Type: replace Abstract: As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated.
By Tomer Keren, Nitay Calderon, Asaf Yehudai, Yotam Perlitz, Michal Shmueli-Scheuer, Roi Reichart
arXiv:2606. 15834v1 Announce Type: new Abstract: The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite systems.
By Yajie Zhou, Ao Li, Ashwin Silla, Zaoxing Liu, Vyas Sekar
arXiv:2602.13217v2 Announce Type: replace
Abstract: Reasoning benchmarks need renewal along two axes: freshness and headroom. VeRA makes both executable and auditable by turning each item into a task...
By Zerui Cheng, Jiashuo Liu, Chunjie Wu, Jiayang Sun, Jianzhu Yao, Pramod Viswanath, Ge Zhang, Wenhao Huang
arXiv:2609. 18366v1 Announce Type: new Abstract: Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed target agent.
By Guojun Zhu, Xunheng Huang, Peng Yin, Jiahui Xie, Sanguo Zhang, Doudou Zhou