arXiv AI

A Descriptive and Normative Theory of Human Beliefs in RLHF

arXiv:2506. 01692v2 Announce Type: replace Abstract: Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.

arXiv Machine Learning
1d ago

The Ethical Decision Head: Operationalizing Normative Ethics in Autonomous Vehicles via Reinforcement Learning from Human Feedback

arXiv:2608. 16710v1 Announce Type: new Abstract: As autonomous vehicles (AVs) approach Level 4 and Level 5 operational capability [SAE International, 2018], their on- board decision systems must handle not only safety-critical locomotion but also their subsequent moral weight.

By Thomas Mbrice, Ammar Ali, Sami Mian, Khai Hern Low, Eric Chen, Arshia Aghajani, Wolf Sch\"afer, Amin Shirangi