arXiv:2608. 01556v1 Announce Type: new Abstract: Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized.
By Seongyoon Kim, Boryeong Cho, Jihwan Oh, Seokhyun Chung, Se-Young Yun
The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.
By Taehyung Kim, Jongeun Choi
The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.
arXiv:2606. 18606v1 Announce Type: cross Abstract: It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable to each community.
By Minsik Oh, Advit Deepak, Sophie Wu, Douwe Kiela, Ekaterina Shutova
The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.
By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao