arXiv AI

KTO: Model Alignment as Prospect Theoretic Optimization

The paper introduces KTO, a new loss function that aligns large language models with human preferences by directly maximizing a prospect-theoretic utility function rather than log-likelihood of preferences. It builds on the idea that existing alignment objectives implicitly use human-aware losses (HALOs) that reflect biases such as loss aversion. KTO achieves performance comparable to or better than preference-based methods across 1B to 30B parameter models while only requiring a binary desirability signal, and it highlights that no single HALO is universally best, emphasizing the importance of choosing the right inductive biases for a given setting.

arXiv AI
Sep 17

Learning Heterogeneous Preferences

The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.

By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk
arXiv AI
3d ago

On the Complexity of Preference-Based Bandits

The paper investigates preference-based bandits where a learner selects pairs of arms and receives binary preference feedback modeled by Bradley–Terry. It introduces the locally sensitive eluder dimension, a new complexity measure for logistic preference feedback, and proposes the GINOP algorithm that uses log-loss confidence sets to balance optimism and exploration. The authors prove a first-order regret bound showing that learning with preference feedback can be as statistically efficient as learning from direct rewards, and they validate their theory with empirical experiments.

By Ahmed Ben Yahmed (CREST, ENSAE Paris, FAIRPLAY), Marc Abeille (FAIRPLAY), Cl\'ement Calauz\`enes (FAIRPLAY)