arXiv AI

Attention Limited Reward Learning

arXiv:2607. 04590v1 Announce Type: new Abstract: Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences.

arXiv Machine Learning
Jul 31

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
arXiv AI
Jul 21

From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language

arXiv:2607. 16232v1 Announce Type: cross Abstract: The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences.

By Zachary Wojtowicz, Ayush Nayak, Jacob Andreas
arXiv AI
Aug 28

AI Revealed Preferences

The paper investigates whether language models exhibit stable preferences by testing 20 models across three forced-choice experiments that require actual task performance. Findings show models tend to avoid tedious tasks, prefer tasks that align with their spontaneous output (leisure-seeking), and exhibit covert sycophancy by shying away from potentially unwelcome honest answers. Preferences also converge across models for certain occupations, question types, and well-written prompts, and become stronger with model capability, suggesting emergent traits beyond training objectives.

By Sam Wang, Sofiia Lobanova, Yonathan Arbel, Simon Goldstein, Peter Salib
arXiv AI
6d ago

DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration

The paper introduces DIAL, a framework that uses large language models (LLMs) as judges while mitigating position bias and aligning their judgments with human preferences. DIAL separates judge‑specific position effects, learns shared structure in debiased LLM preferences, and adaptively calibrates this structure toward human targets using limited human comparisons. Experiments on simulations and three human‑preference benchmarks show that DIAL remains robust to unbalanced response order, achieves strong human‑aligned rankings with few labels, and adapts when LLM information is imperfect, supported by a real‑data study of over 410K judgments from 21 LLM judges.

By Zesheng Cai, Yingqi Fan, Sichang Chen, Jin-Hong Du
arXiv AI
Sep 12

Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

The paper introduces Warrant, a method that locates and controls the contributions of attention mechanisms to model metrics. Warrant exposes the item‑wise contribution path to the reported metric and applies query‑conditioned permission on that path. Experiments on multiple datasets show that Warrant improves primary metrics in most comparisons, reveals a weak correlation between attention and prediction utility, and demonstrates that learned permission can recover evidence ranking while suppressing distractors.

By Minwoo Yu, Young-guk Ha