arXiv Machine Learning By Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha

When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO

Read the original on arXiv Machine Learning →

The paper introduces Guidance-Augmented GRPO (GA‑GRPO), a theoretical framework that unifies several external‑guidance methods for reinforcement learning with verifiable rewards (RLVR) used to improve large language model reasoning. GA‑GRPO treats external guidance as a stochastic operator that rewrites the question distribution, yielding a biased on‑policy policy‑gradient estimator whose bias is bounded by the total‑variation guidance divergence. The authors prove convergence rates, derive an optimal guidance‑weight formula, and validate their predictions experimentally on a math‑reasoning benchmark, showing that the optimal‑weight GA‑GRPO outperforms existing methods while saving GPU time.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 16

A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions

arXiv:2606. 16733v1 Announce Type: new Abstract: Policy gradient algorithms for language models optimize the same objective $J(\theta) = \mathbb{E}*{\tau \sim p*\theta(\tau)}[R(\tau)]$, which has exactly two factors: the trajectory probability $p_\theta(\tau)$ and the reward $R(\tau)$.

By Jianghan Shen, Siqi Luo, Yue Li, Jiyao Liu, Wanying Qu, Yi Zhang, Ziyan Huang, Tianbin Li, Ming Hu, Xiaohong Liu, Yirong Chen, Junjun He
arXiv Machine Learning
Jul 31

Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL Finetuning

arXiv:2607. 27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critically depends on selecting prompts of appropriate difficulty for the current policy.

By Haodong Zhu, Yangyang Ren, Yanjing Li, Sheng Xu, Haiguang Liu, Linlin Yang, Baochang Zhang
arXiv Machine Learning
Jun 11

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

arXiv:2606. 11709v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distribution it produces under privileged context, typically a verified solution.

By Leyi Pan, Shuchang Tao, Yunpeng Zhai, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, Lijie Wen
arXiv Machine Learning
5d ago

Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients

The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.

By Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang