arXiv AI By Sylee Dandekar, Shripad Deshmukh, Frank Chiu, W. Bradley Knox, Scott Niekum

A Descriptive and Normative Theory of Human Beliefs in RLHF

Read the original on arXiv AI →

arXiv:2506. 01692v2 Announce Type: replace Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.