arXiv AI

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

arXiv:2605. 07724v2 Announce Type: replace-cross Abstract: Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed reward signal, the model tends to collapse onto a narrow set of outputs that over-optimize that objective.

Hugging Face Trending Papers
Aug 19

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that adapts to diverse user reward models in AI alignment. By automatically building a curriculum that branches and reuses reward models, it addresses the problem of users whose reward models are hard to optimize, a group often underserved by conventional methods. Experiments on personalized continuous control demonstrate that CurriPO improves population satisfaction by 1.2–2.1× over the best baseline while cutting training time.

arXiv Machine Learning
Aug 20

To Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization

The paper introduces CurriPO, a tree‑structured curriculum that automatically adapts to diverse user reward models in AI alignment tasks. By exploiting the natural hierarchy between easy‑ and hard‑to‑optimize reward models, CurriPO covers a broad user population in a single traversal, reusing previously incorporated reward models. Experiments on personalized continuous control show that CurriPO improves population satisfaction by 1.2–2.1× over the strongest baseline while cutting training time and better serving users traditionally underserved by conventional optimization.

By Taehyung Kim, Jongeun Choi
arXiv Machine Learning
Jul 3

Optimizing Visual Generative Models via Distribution-wise Rewards

arXiv:2607. 02291v1 Announce Type: new Abstract: Conventional reinforcement learning strategies for visual generation typically employ sample-wise reward functions, yet this practice frequently results in reward hacking that degrades image diversity and introduces visual anomalies.

By Ruihang Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang
arXiv AI
Sep 1

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.

By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv AI
Aug 6

Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation

arXiv:2601. 12401v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks.

By Jinmei Liu, Haoru Li, Zhenhong Sun, Chaofeng Chen, Yatao Bian, Bo Wang, Daoyi Dong, Chunlin Chen, Zhi Wang