arXiv Machine Learning
Sep 2

Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?

The study investigates whether small differences in leaderboard rankings between large language models (LLMs) are robust to changes in benchmark composition. Using item‑level responses from five benchmarks and a spectral approximation to multidimensional item‑response theory, the authors find that while overall rankings remain highly correlated, a significant portion (30.9–47.1%) of near‑tie pairs reverse order when benchmark items are recomposed based on low differential item functioning. This suggests that sub‑one‑percentage‑point leaderboard gaps may not reliably reflect true model superiority.

By Qiaoyuan Zheng, Yiqu Yang