arXiv Machine Learning By Chen Yang, Xianyang Zhang, Jun Chen

How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?

Read the original on arXiv Machine Learning →

The paper investigates how many hidden model variants can exist while still supporting a published leaderboard margin that shows a provider’s advantage over a fixed comparator. It derives a sensitivity curve for a fixed candidate family under a Gaussian margin model, linking the maximum number of hidden variants to a lower bound on within‑family correlation. Using this framework, an audit of 394 adjacent‑rank claims on the Open LLM Leaderboard found that 391 lack statistical support before any correction, and that certification of the remaining claims depends on assumptions about the hidden family’s correlation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 2

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

The study investigates whether small differences in leaderboard rankings between large language models (LLMs) are robust to changes in benchmark composition. Using item‑level responses from five benchmarks and a spectral approximation to multidimensional item‑response theory, the authors find that while overall rankings remain highly correlated, a significant portion (30.9–47.1%) of near‑tie pairs reverse order when benchmark items are recomposed based on low differential item functioning. This suggests that sub‑one‑percentage‑point leaderboard gaps may not reliably reflect true model superiority.

By Qiaoyuan Zheng, Yiqu Yang
arXiv Computation and Language
Sep 4

Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards

The paper investigates how benchmark contamination—leakage of test items into training data—affects large language model (LLM) leaderboards. By comparing original test items with semantically equivalent paraphrases, the authors measure contamination as a violation of anchor-item invariance and find that it inflates absolute scores but rarely changes model rankings. Across 47 public models and 74 finetuned models on four benchmarks, the rank correlation between standard and paraphrase-controlled leaderboards is 0.997, with only a handful of cases showing differential contamination that could alter rankings.

By Xingyao Xiao (Stanford University), Yihong Cheng (City University of Macau)
arXiv AI
6d ago

Accounting for Bias Enables Sustainable LLM Evaluation

The paper argues that the current LLM-as-a-judge evaluation method, which compensates for systematic measurement bias by increasing the number of comparisons, is statistically unsound and computationally wasteful. It identifies that treating LLM judges as neutral ignores documented biases such as position bias, verbosity bias, judge severity, and self‑enhancement. The authors propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, enabling reliable rankings with far fewer comparisons and negligible additional compute.

By Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer