arXiv Machine Learning By Seongyoon Kim, Boryeong Cho, Jihwan Oh, Seokhyun Chung, Se-Young Yun

Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning

Read the original on arXiv Machine Learning →

arXiv:2608. 01556v1 Announce Type: new Abstract: Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 18

Global Federated Learning Strategies for Building Efficient Personalized Models

arXiv:2608. 15107v1 Announce Type: new Abstract: Federated learning (FL) is a practical framework that can train models on distributed user data while guaranteeing data privacy; however, due to heterogeneity in which each user has a different data distribution, problems frequently arise where both global and personalization performance deteriorate simultaneously.

By Seongyoon Kim
arXiv AI
Sep 1

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.

By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv Machine Learning
1d ago

From Task Mixtures to Specialized Experts

The paper investigates federated learning where each client’s data consists of unknown mixtures of distinct tasks, a scenario termed compound heterogeneity. It shows that when tasks share a common feature geometry, the optimal model for a mixed client is a convex combination of task‑specific models, motivating input‑dependent routing to specialized experts. The authors propose FedSEE, a method that recovers task experts via a convex program and achieves better performance than baselines, reducing negative transfer by 2.9 points overall and 3.7 points for the worst‑served quartile.

By Hojat Allah Salehi, Mehrdad Mahdavi, Andrew Arash Mahyari, M. Hadi Amini
arXiv Machine Learning
Aug 20

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.

By Taehyung Kim, Jongeun Choi