arXiv Machine Learning

A Control Function Framework for Mitigating Position Bias in Learning to Rank Systems

arXiv:2506. 06989v3 Announce Type: replace-cross Abstract: Learning-to-rank (LTR) systems commonly depend on implicit feedback, such as user clicks, because it is easy to collect and can serve as a valuable signal of user preferences.

arXiv Machine Learning
Sep 1

Adaptive Doubly Robust Off-Policy Evaluation for Ranking Policies under Diverse User Behavior

The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.

By Kosuke Iguchi, Ren Kishimoto
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv Statistics ML
Sep 4

Low Rank for Rank: Uncertainty-Aware Task-Specific LLM Ranking under Sparse Pairwise Comparisons

The paper introduces a low‑rank framework for ranking large language models (LLMs) on task‑specific benchmarks using sparse pairwise comparisons. By modeling the task‑by‑model ability matrix as low rank, the method shares information across related tasks while preserving task‑specific differences, and it provides uncertainty‑aware ranking through debiased estimators and simultaneous confidence sets. Experiments on synthetic data and the Chatbot Arena benchmark demonstrate improved sample efficiency and tighter, better‑calibrated ranking certificates, especially in the sparse comparison regime typical of real LLM evaluations.

By Jiachun Li, David Simchi-Levi, Will Wei Sun