arXiv Machine Learning
1d ago

Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients

The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.

By Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang