arXiv AI

Can LLMs Rank? A Tale of Triads and Triage

arXiv:2606. 30412v1 Announce Type: cross Abstract: From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered as judges of consequential decisions that require ranking people for scarce resources.

arXiv Computation and Language
4d ago

Pair Difficulty Matters: Rethinking Pairwise LLM-as-a-Judge Evaluation and Consistency

Large Language Model judges are commonly used to rank texts via pairwise comparison, with reliability traditionally measured by position bias, transitivity, and pairwise agreement. This paper argues that these proxies are misleading because they are dominated by close‑rank‑gap pairs, which contribute little to the overall ranking, while far‑gap pairs carry the true ranking signal. Experiments on simulations and human‑rated corpora show weak correlation between the proxies and actual ranking accuracy, suggesting judges should be evaluated using rank‑gap‑conditional metrics against human rankings.

By Bruno Brocai, Maria Becker
arXiv AI
2d ago

Agent Evaluation Reliability: More Tasks Won't (Always) Fix An Agent Leaderboard

Agent evaluations increasingly benchmark LLMs, but rankings can be swayed by evaluation conditions such as scaffolds or tasks, making reliability claim‑dependent. A Bayesian variance‑decomposition framework applied to 22 benchmarks shows that reliability varies with the measurement goal: fixed model‑scaffold systems rank reliably, while underlying‑model rankings are less stable. Scaffold choice can alter conclusions, and adding more tasks only modestly improves reliability when scaffold coverage is limited; however, pooling diverse benchmarks can substantially raise cross‑task ranking reliability and reduce cost.

By Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
Hugging Face Trending Papers
Aug 3

Aggregate-then-Calibrate for Human-centered Assessment with Theoretical Guarantees

Human-centered assessment tasks, which are essential for systematic decision-making, rely heavily on human judgment and typically lack verifiable ground truth. Existing approaches face a dilemma: methods using only human judgments suffer from heterogeneous expertise and inconsistent rating scales, while methods using only model-generated scores must learn from imperfect proxies or incomplete features.

arXiv Machine Learning
Sep 14

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

GAUGE is a new offline protocol that evaluates whether the common practice of using an LLM-as-a-judge to rank task‑oriented agents actually aligns with a verifiable reward. Across 25 agents from six providers on two benchmarks, GAUGE finds that user satisfaction scores are largely uncorrelated with task success, and that the judge’s ranking loses precision when agents are closely matched in performance. The study highlights a gap between ranking validity and construct validity in current evaluation practices.

By Umesh Bodhwani, Thanh Tran, Kai Wei
arXiv AI
6d ago

Accounting for Bias Enables Sustainable LLM Evaluation

The paper argues that the current LLM-as-a-judge evaluation method, which compensates for systematic measurement bias by increasing the number of comparisons, is statistically unsound and computationally wasteful. It identifies that treating LLM judges as neutral ignores documented biases such as position bias, verbosity bias, judge severity, and self‑enhancement. The authors propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, enabling reliable rankings with far fewer comparisons and negligible additional compute.

By Harshita Katoch, David Antony Selby, Gerrit Gro{\ss}mann, Sebastian Vollmer
arXiv AI
Aug 19

LLM-Derived Preference Judgments Are Not Self-Consistent

The paper investigates whether large language models (LLMs) produce self‑consistent numerical preference judgments when asked to estimate a person’s willingness to pay for items. By comparing differences in stated willingness to pay with the payment that would make a person indifferent between items, the authors develop statistical tests to measure deviations from a single utility function. Experiments on flight, apartment, and hotel scenarios across six LLMs show persistent inconsistencies, indicating that LLM‑derived preference judgments cannot be reliably summarized by one utility function.

By Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
arXiv Machine Learning
Aug 28

Equal Ranking Quality, Different Decisions: Training Order-Consistent LLM Scorers

The paper investigates how the order of candidate documents in a prompt affects the decisions made by large‑language‑model (LLM) scorers, even when their ranking quality is similar. It shows that five scorers with only a 0.010 nDCG@10 difference can produce retained‑set overlaps as low as 0.66–0.84, and that existing rerankers still exhibit significant order dependence. The authors propose Order‑Consistency SFT (OC‑SFT), a training method that reduces this dependence, maintaining ranking quality while improving decision stability across multiple tasks.

By Markus Frohmann, Mahdiyar Alavi, Elizabeth Lingg, Navid Rekabsaz
arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand