arXiv AI By Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, Chuangchuang Wang

Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

Read the original on arXiv AI →

The paper introduces FastRL, a reinforcement learning framework designed to enhance the efficiency of Group Relative Policy Optimization (GRPO) and its variants. FastRL employs an advantage-aware pruning strategy that retains high-advantage trajectories while preserving gradient diversity, and an adaptive rollout sampling mechanism that adjusts sampling scale during training based on historical pruning data. Experiments show that FastRL can be integrated into GRPO, DAPO, and GSPO, yielding a 2.07× speedup on Geometry3K and GeoQA8K-R1V and a 1.64% accuracy improvement on visual reasoning benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jun 16

DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning

arXiv:2505. 09655v5 Announce Type: replace-cross Abstract: Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning.

By Xiwen Chen, Wenhui Zhu, Peijie Qiu, Xuanzhao Dong, Hao Wang, Haiyu Wu, Huayu Li, Aristeidis Sotiras, Yalin Wang, Abolfazl Razi