arXiv Machine Learning By Simone Drago, Marco Mussi, Leonardo Bianconi, Alberto Maria Metelli

Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

Read the original on arXiv Machine Learning →

arXiv:2607. 11432v1 Announce Type: new Abstract: In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.