arXiv Machine Learning

Multiple latent orderings better predict language model preferences

The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.

arXiv AI
Jun 3

Distribution-Calibrated Inference Time Compute for Thinking LLM-as-a-Judge

arXiv:2512. 03019v2 Announce Type: replace-cross Abstract: Thinking Large Language Models (LLMs) used as judges for pairwise preferences remain noisy at the single-sample level, and common aggregation rules (majority vote, soft self-consistency, or instruction-based self-aggregation) are inconsistent when ties are allowed.

By Hamid Dadkhahi, Firas Trabelsi, Parker Riley, Juraj Juraska, Mehdi Mirzazadeh
arXiv AI
Jul 21

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.

By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
arXiv AI
Aug 19

LLM-Derived Preference Judgments Are Not Self-Consistent

The paper investigates whether large language models (LLMs) produce self‑consistent numerical preference judgments when asked to estimate a person’s willingness to pay for items. By comparing differences in stated willingness to pay with the payment that would make a person indifferent between items, the authors develop statistical tests to measure deviations from a single utility function. Experiments on flight, apartment, and hotel scenarios across six LLMs show persistent inconsistencies, indicating that LLM‑derived preference judgments cannot be reliably summarized by one utility function.

By Matthew T. Ford, Francis Bahk, Jingjing Wang, Adam S. Jovine, Tinghan Ye, David B. Shmoys, Peter I. Frazier
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv Computation and Language
Aug 31

Auditing LLM Benchmarks with Item Response Theory

The paper introduces an Item Response Theory (IRT)–based indicator that identifies likely mislabeled items in large language model (LLM) benchmarks with 95% precision among the top 200 examples across seven preference and multiple-choice datasets, using responses from 114 models. It outperforms a supervised classifier and attributes the mislabels to mechanical labeling heuristics, inherited annotation errors, and inherently ambiguous items. The IRT analysis also reveals that reward models tend to specialize in stylistic preference rather than factual knowledge, and pinpoints a frontier reward model that aligns with detected mislabels at 78% accuracy compared to 38% for other models, suggesting benchmark contamination or over‑optimization.

By Sander Land, Daniel M. Bikel
arXiv AI
2d ago

Evaluating LLM-Generated Preference Distributions

The paper evaluates how Large Language Models generate preference distributions for air travel, restaurants, and consumer products. It finds that while each model produces self-coherent outcomes that stabilize quickly, there is significant disagreement across different model families and scales, with little consensus even on the most probable preferences. These discrepancies persist across various decoding strategies, temperature settings, and prompt variations, indicating that the model choice itself has a larger impact than prompt wording.

By Fan Huang, Minsuk Kim, C. Tyler Diggans, Filippo Radicchi