arXiv Machine Learning

MO-GRPO: Mitigating Reward Hacking of Group Relative Policy Optimization on Multi-Objective Problems

arXiv:2509. 22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available.

arXiv Machine Learning
Jun 5

MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following

arXiv:2606. 06058v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is ideal for multi-constraint instruction following, yet standard group-relative policy optimization (GRPO) becomes unstable under discrete, low-dispersion rewards, where within-group reward distributions are frequently homogeneous.

By Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili, Mohammad Javad Dousti
arXiv Machine Learning
Jun 16

GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization

arXiv:2606. 16771v1 Announce Type: new Abstract: As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities.

By Haotian Liu, Yihao Liu, Jingwei Ni, Siyuan Huang, Xinpeng Liu, Pengyu Cheng, Jiajun Song, Ruijin Ding, Junfeng Li, Zhechao Yu, Mengyu Zhou, Hongteng Xu, Xiaoxi Jiang, Guanjun Jiang
arXiv AI
Sep 1

Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment

The paper introduces Personalized Group Relative Policy Optimization (P‑GRPO), a new alignment framework for large language models that separates advantage estimation from immediate batch statistics. By normalizing advantages using preference‑group‑specific reward histories instead of the concurrent generation group, P‑GRPO maintains contrastive signals for distinct user preferences. Experiments across various tasks show that P‑GRPO converges faster and yields higher rewards than standard GRPO, improving alignment with heterogeneous human preferences while preserving general capabilities.

By Jialu Wang, Heinrich Peters, Asad A. Butt, Navid Hashemi, Alireza Hashemi, Pouya M. Ghari, Joseph Hoover, James Rae, Morteza Dehghani
arXiv Machine Learning
1d ago

Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL

The paper introduces Density-Aware Reward Aggregation (DARA), a method that adjusts reward weighting in multi-reward reinforcement learning based on the density of active rewards within rollout batches. By applying an inverse-square-root density correction, DARA gives more weight to less frequently active rewards, enabling faster learning of targeted behaviors. Experiments on tool calling and mathematical reasoning demonstrate that DARA achieves comparable final performance while reducing training steps by up to 26% and 65% respectively.

By Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu
Hugging Face Trending Papers
Sep 3

Spurious Advantage Hidden in GRPO

The paper identifies a flaw in Group Relative Policy Optimization (GRPO) called the spurious advantage, where the advantage estimator can incorrectly reward rollouts that reach correct answers by guessing rather than reasoning. This issue appears in bounded-answer tasks, open-answer sets with bounded sub-cases, and search agents with many paths to the same answer. The authors propose SIGNBALANCE, a new magnitude estimator that removes composition, maintains verifier sign, applies a global scale, and uses stop‑gradient per‑class rescaling to restore zero‑mean balance, achieving comparable or better performance on math and search benchmarks.