Accounting for Bias Enables Sustainable LLM Evaluation
Read the original on arXiv AI →The paper argues that the current LLM-as-a-judge evaluation method, which compensates for systematic measurement bias by increasing the number of comparisons, is statistically unsound and computationally wasteful. It identifies that treating LLM judges as neutral ignores documented biases such as position bias, verbosity bias, judge severity, and self‑enhancement. The authors propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, enabling reliable rankings with far fewer comparisons and negligible additional compute.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.