arXiv Machine Learning By Qiaoyuan Zheng, Yiqu Yang

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

Read the original on arXiv Machine Learning →

The study investigates whether small differences in leaderboard rankings between large language models (LLMs) are robust to changes in benchmark composition. Using item‑level responses from five benchmarks and a spectral approximation to multidimensional item‑response theory, the authors find that while overall rankings remain highly correlated, a significant portion (30.9–47.1%) of near‑tie pairs reverse order when benchmark items are recomposed based on low differential item functioning. This suggests that sub‑one‑percentage‑point leaderboard gaps may not reliably reflect true model superiority.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand
arXiv Machine Learning
Aug 28

Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers

The paper investigates how the order of candidate documents in a prompt affects the decisions made by large‑language‑model (LLM) scorers, even when their ranking quality is similar. It shows that five scorers with only a 0.010 nDCG@10 difference can produce retained‑set overlaps as low as 0.66–0.84, and that existing rerankers still exhibit significant order dependence. The authors propose Order‑Consistency SFT (OC‑SFT), a training method that reduces this dependence, maintaining ranking quality while improving decision stability across multiple tasks.

By Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz
arXiv AI
Jun 4

Knowledge Index of Noah's Ark

arXiv:2606. 05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets.

By Sheng Jin, Minghao Liu, Yunze Xiao, Zeqi Zhou, Heli Qi, Yifan Yao, Meishu Song, Kaijing Ma, Xuan Zhang, Sicong Jiang, Yizhe Li, Ningshan Ma, Jie Wei, Ziniu Li, Minglai Yang, Bangya Liu, Yiming Liang, Xiao Fang, Qingcheng Zeng, Jiarui Liu, Rui Yang, Shen Yan, Wenhao Huang, Jiaheng Liu, Zihan Wang, Weihao Xuan, Ge Zhang