Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency
Read the original on arXiv Computation and Language →Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.