arXiv AI

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis

arXiv Machine Learning
1d ago

Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients

The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.

By Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang
Hugging Face Trending Papers
Jul 21

Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning

Asynchronous reinforcement learning improves throughput by decoupling rollout generation from optimization, but staleness is an inevitable byproduct compounded by policy lag, engine delays, and mixture-of-experts routing. From a trust-region perspective, this mismatch is critical: training-inference divergence governs approximation error in finite-horizon bounds, whereas PPO clipping only gates sampled outward updates, acting as a sampled surrogate rather than a full-policy constraint.

arXiv Machine Learning
Sep 2

Group Adaptive Clipping Policy Optimization

Group Adaptive Clipping Policy Optimization (GAPO) is a plug‑in modification to GRPO methods that adapts the importance‑sampling clipping boundary based on rollout advantage. By allowing rollouts with larger learning signals to receive proportionally greater update headroom, GAPO addresses the limitation of fixed clipping that suppresses rare but informative rollouts. Experiments on Qwen and Llama models show that GAPO consistently improves Pass@1 and Pass@k on math reasoning and coding benchmarks where base model pass rates are low.

By Sheng Jia, Xiao Wang, Shiva Prasad Kasiviswanathan, Rein Houthooft
arXiv AI
Jun 9

Sparrow: Sparse Rollout for Stable and Efficient Long-context RL of Large Language Models

arXiv:2606. 08446v1 Announce Type: cross Abstract: Despite being powerful, reinforcement learning with verifiable rewards (RLVR) induces extremely long COT, making it computationally expensive.

By Yang Zhou, Ranajoy Sadhukhan, Zhaofeng Sun, Zhuoming Chen, Souvik Kundu, Saket Dingliwal, Sai Muralidhar Jayanthi, Aram Galstyan, Haizhong Zheng, Beidi Chen
Hugging Face Trending Papers
Aug 20

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods estimate policy advantages from multiple rollouts, avoiding the substantial memory overhead of conventional proximal policy optimization (PPO) and achieving strong performance on long-horizon interactive tasks.

arXiv AI
Aug 3

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

arXiv:2607. 22186v2 Announce Type: replace Abstract: Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse.

By Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen