arXiv Machine Learning By Brahim Driss, Alex Davey, Riad Akrour

PB$^2$: Preference Space Exploration via Population-Based Methods in Preference-Based Reinforcement Learning

Read the original on arXiv Machine Learning →

arXiv:2506. 13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 19

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.

arXiv Machine Learning
Aug 20

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.

By Taehyung Kim, Jongeun Choi