CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2509. 22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available.
arXiv:2609.00213v1 Announce Type: new Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific e...
The paper introduces Density-Aware Reward Aggregation (DARA), a method that adjusts reward weighting in multi-reward reinforcement learning based on the density of active rewards within rollout batches. By applying an inverse-square-root density correction, DARA gives more weight to less frequently active rewards, enabling faster learning of targeted behaviors. Experiments on tool calling and mathematical reasoning demonstrate that DARA achieves comparable final performance while reducing training steps by up to 26% and 65% respectively.
The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.
arXiv:2606. 26300v1 Announce Type: new Abstract: A classical intuition holds that verifying a solution is easier than producing one.
Range-GRPO introduces a semi‑supervised post‑training framework that uses conformally calibrated reward ranges instead of single point scores for large language models. By comparing reward ranges pairwise within rollout groups, the method incorporates reward uncertainty into both the magnitude and direction of learning signals. Experiments show that Range‑GRPO outperforms other semi‑supervised approaches on both in‑distribution and out‑of‑distribution tasks while using fewer training resources.