arXiv AI

A Descriptive and Normative Theory of Human Beliefs in RLHF

arXiv:2506. 01692v2 Announce Type: replace Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.

arXiv AI
Aug 20

How AI Prompts Can Teach Us About the Structure of Human Behavior

The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.

By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
arXiv AI
Sep 17

Learning Heterogeneous Preferences

The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.

By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk
arXiv Machine Learning
Aug 18

The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

arXiv:2608. 16710v1 Announce Type: new Abstract: As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight.

By Thomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low, Eric Chen, Arshia Aghajani, Wolf Sch\"afer, Amin Shirangi
arXiv AI
Sep 15

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.

By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez