arXiv AI By Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum

A Descriptive and Normative Theory of Human Beliefs in RLHF

Read the original on arXiv AI →

arXiv:2506. 01692v2 Announce Type: replace Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 20

How AI Prompts Can Teach Us About the Structure of Human Behavior

The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.

By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv AI
Sep 17

Learning Heterogeneous Preferences

The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.

By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk