arXiv Machine Learning

On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

arXiv:2607. 23364v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1.

arXiv Machine Learning
1d ago

Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients

The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.

By Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang
arXiv AI
Sep 25

To Think or Not to Think: Allocating Reasoning Where It Helps

The paper introduces CARE, a contrastive accuracy reward estimation method that adaptively adjusts reasoning length for large language models. By comparing beneficial length adjustments from online sampled responses, CARE applies adaptive length rewards within Group Relative Policy Optimization without extra hyperparameters or inference cost. Experiments on multiple reasoning benchmarks show that CARE improves Pass@1 by up to 4% while reducing reasoning length by 37%, achieving higher token efficiency.

By Zhengdong He, Yunfan Zhou, Jianguo Yao, Haibing Guan, Xijun Li
arXiv Machine Learning
Jun 5

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

arXiv:2606. 05434v1 Announce Type: new Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks, but it treats every token position and every sampled rollout symmetrically.

By Chirag Chawla, Rohan Charudatt Salvi, Madhav S. Baidya