arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.
By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh
arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.
By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.
By Yujun Zhou, Christopher M. Ackerman
The paper investigates whether large language models (LLMs) produce self‑consistent numerical preference judgments when asked to estimate a person’s willingness to pay for items. By comparing differences in stated willingness to pay with the payment that would make a person indifferent between items, the authors develop statistical tests to measure deviations from a single utility function. Experiments on flight, apartment, and hotel scenarios across six LLMs show persistent inconsistencies, indicating that LLM‑derived preference judgments cannot be reliably summarized by one utility function.
By Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
arXiv:2602.02219v3 Announce Type: replace
Abstract: Large language models are widely employed as evaluators, a paradigm commonly referred to as LLM-as-a-judge. Prior research has predominantly examin...
By Yuzheng Xu, Tosho Hirasawa, Tadashi Kozuno, Yoshitaka Ushiku
The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.
By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv:2506. 14092v4 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes domains such as hiring and university admissions, where choices often involve selecting among competing alternatives.
By Haonan Yin, Shai Vardi, Vidyanand Choudhary
arXiv:2602. 10286v3 Announce Type: replace Abstract: Pairwise preference learning is central to machine learning, with recent applications in aligning language models with human preferences.
By Rattana Pukdee, Maria-Florina Balcan, Pradeep Ravikumar
arXiv:2607. 02672v1 Announce Type: new Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.
By Bailey Flanigan, Michelle Si
The paper introduces an Item Response Theory (IRT)–based indicator that identifies likely mislabeled items in large language model (LLM) benchmarks with 95% precision among the top 200 examples across seven preference and multiple-choice datasets, using responses from 114 models. It outperforms a supervised classifier and attributes the mislabels to mechanical labeling heuristics, inherited annotation errors, and inherently ambiguous items. The IRT analysis also reveals that reward models tend to specialize in stylistic preference rather than factual knowledge, and pinpoints a frontier reward model that aligns with detected mislabels at 78% accuracy compared to 38% for other models, suggesting benchmark contamination or over‑optimization.
By Sander Land, Daniel M. Bikel
The paper evaluates how Large Language Models generate preference distributions for air travel, restaurants, and consumer products. It finds that while each model produces self-coherent outcomes that stabilize quickly, there is significant disagreement across different model families and scales, with little consensus even on the most probable preferences. These discrepancies persist across various decoding strategies, temperature settings, and prompt variations, indicating that the model choice itself has a larger impact than prompt wording.
By Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi
arXiv:2608. 11947v1 Announce Type: cross Abstract: Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option order, which makes them unreliable measures of model knowledge.
By Karl Hanna, Chen Feng