arXiv Computation and Language By Shachar Don-Yehiya, Asaf Yehudai, Leshem Choshen, Omri Abend

Mediocrity is the key for LLM as a Judge Anchor Selection

Read the original on arXiv Computation and Language →

The paper examines how the choice of anchor model in LLM-as-a-judge evaluations affects reliability. By testing 22 anchors on the Arena-Hard-v2.0 dataset, it shows that extreme anchors (best or worst performers) are poor choices, reducing correlation with human rankings. The study quantifies the anchor effect size, compares it to judge model selection, and offers guidelines and power‑analysis recommendations for more reliable benchmark design.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 3

PACE: A Proxy for Agentic Capability Evaluation

arXiv:2607. 02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure.

By Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng, Yunze Xiao, Daniel Lee, Aditya Bharat Soni, Vincent Lo, Xiang Yue, Graham Neubig
arXiv AI
Jun 4

Knowledge Index of Noah's Ark

arXiv:2606. 05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets.

By Sheng Jin, Minghao Liu, Yunze Xiao, Zeqi Zhou, Heli Qi, Yifan Yao, Meishu Song, Kaijing Ma, Xuan Zhang, Sicong Jiang, Yizhe Li, Ningshan Ma, Jie Wei, Ziniu Li, Minglai Yang, Bangya Liu, Yiming Liang, Xiao Fang, Qingcheng Zeng, Jiarui Liu, Rui Yang, Shen Yan, Wenhao Huang, Jiaheng Liu, Zihan Wang, Weihao Xuan, Ge Zhang