arXiv AI

Internal Pluralism and the Limits of Pairwise Comparisons

arXiv:2607. 02672v1 Announce Type: new Abstract: Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.

arXiv AI
Jul 21

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.

By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
arXiv AI
4d ago

JudgeProfile: Understanding and Steering Subjectivity in LLM Judges

JudgeProfile is a framework that analyzes the subjectivity of large language model (LLM) judges by separating evaluation into perception—how judges compare responses on attributes such as clarity, correctness, and detail—and prioritization—how much each attribute influences the final decision. Using the curated SubjectiveSet dataset of 50,013 response pairs evaluated by 21 judges across 87 attributes, the study finds that judges often agree on attribute judgments even when their overall choices differ. By estimating and adjusting attribute weights, the authors improve agreement with reference labels from 66.48% to 71.97%, outperforming fine‑tuning and rubric prompting.

By Qi Cao, Kangning Liu, Xuan Kan, Shunwen Tan, Yang Pei, Dake Chen, Yatai Ji, Zixuan Ye, Yuanpeng Tu, Daniel Li, Junbiao Tang, Pengtao Xie, Zihao He
arXiv Machine Learning
Sep 22

Multiple latent orderings better predict language model preferences

The paper argues that language models’ intransitive preferences arise from multiple internally consistent latent orderings rather than noise around a single ordering. By demonstrating that a single ordering cannot explain observed inconsistencies and introducing a noise‑augmented mixture Bradley‑Terry model, the authors show that mixtures of orderings better capture preference structure across several models and tasks. A case study on Moral Machine dilemmas further illustrates that models can share latent components even when aggregate preferences differ.

By Aviral Chawla, William H. W. Thompson, Jean-Gabriel Young
arXiv AI
Sep 4

Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy

The paper examines how Preference Inference (PI) models used in large-scale participatory democracy platforms can alter the perceived consensus and minority support by predicting missing votes. It introduces a collective‑centric evaluation framework that assesses whether inferred votes maintain key properties of the overall preference landscape, rather than focusing solely on individual prediction accuracy. Using the largest multilingual dataset to date—four consultations with over 90,000 participants, 1 million votes, and 22 languages—the study finds that models with similar predictive accuracy can differ markedly in how well they preserve the collective structure, underscoring that accuracy alone is insufficient for evaluating PI in democratic contexts.

By Pierre-Antoine Lequeu, Salim Hafid, Paul Lerner, Nazanin Shafiabadi, Laur\`ene Cave, David Mas, Jean-Philippe Cointet, Benjamin Piwowarski, Fran\c{c}ois Yvon
arXiv AI
Sep 17

Learning Heterogeneous Preferences

The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.

By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk
arXiv AI
Sep 24

Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment

The paper investigates how to create steerable AI models that can balance multiple, sometimes conflicting objectives, a necessity for pluralistic alignment. Using Multi-Objective Direct Preference Optimization (MODPO), the authors examine when a single model can improve two objectives simultaneously and how to cover many trade‑offs without training separate models. They find that two pre‑training measurements predict objective alignment for human‑annotated data but not for AI‑annotated data, and that selecting the nearest trained model or merging parameters can broaden trade‑off coverage, though neither approach consistently matches direct training.

By David Tsoi, Esra D\"onmez
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du