arXiv:2606. 09124v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals.
By Suhwan Kim, Taehyun Cho, Geon-Hyeong Kim, Yu Jin Kim, Youngsoo Jang, Moontae Lee, Jungwoo Lee
The paper presents an AI-based method that uses a large language model to emulate human decision-making by assigning it a "type vector" describing traits such as Altruism and Risk Aversion. By varying these dimensions and values, the authors fit the model to over 119,000 decisions from 78,657 participants in 10 classic economic games, finding that three dimensions—Risk Aversion, Strategic Sophistication, and Trust—sufficiently capture human behavior. The resulting type clusters, fewer than a dozen, predict behavior in new games, suggesting a low-dimensional, portable representation of human behavior across diverse settings.
By Matthew O. Jackson, Benjamin S. Manning, Yutong Xie, Walter Yuan, Qiaozhu Mei
The paper introduces a method for learning heterogeneous, individually conditioned utility functions—termed individuated utility—by leveraging rational choice theory. It presents a multi-stage architecture that estimates these functions from multimodal data and evaluates it on a large dataset of aesthetic judgments about automotive wheel designs. Results show that individuated models outperform universal utility models and foundation baselines, indicating that annotator disagreement reflects meaningful preference diversity.
By Shiwali Mohan, Matt Hong, Dule Shu, Aniek Fransen, Shabnam Hakimi, Matt Klenk
arXiv:2608. 10126v1 Announce Type: cross Abstract: Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity.
By M P V S Gopinadh, Karthik Kamuju, Kummari Avinash, John Joshua, Srinivasa Raju Rudraraju
arXiv:2606. 30863v1 Announce Type: new Abstract: Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified.
By Irena Saracay, Ludwig Schmidt, Carlos Guestrin
arXiv:2608. 12339v1 Announce Type: cross Abstract: Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases.
By Eldad Yechiam, Adi Tarabeih
arXiv:2607. 16195v1 Announce Type: new Abstract: We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF).
By Elena Kopteva, Vitaliy Hlynianyi-Zhuk
arXiv:2606. 22974v2 Announce Type: replace Abstract: Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure.
By Yujun Zhou, Christopher M. Ackerman
arXiv:2608. 16710v1 Announce Type: new Abstract: As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight.
By Thomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low, Eric Chen, Arshia Aghajani, Wolf Sch\"afer, Amin Shirangi
arXiv:2607. 11632v1 Announce Type: new Abstract: Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality.
By Jiangtao Han, Shoufeng Ma, Shuxian Xu, Geng Li, Shuai Ling, Ning Jia, Zhengbing He
The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.
By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez
arXiv:2606. 17657v1 Announce Type: new Abstract: People make decisions differently in strategic interactions.
By Zirui Cheng, Zeyu Shen, Thomas L. Griffiths, Peter Henderson