arXiv Machine Learning By Junoh Park, Junseo Hwang, Wonguk Cho, Taesup Kim

ReCo: Reweighting GRPO Against Distributional Concentration

Read the original on arXiv Machine Learning →

arXiv:2607. 26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 5

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

arXiv:2606. 05434v1 Announce Type: new Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks, but it treats every token position and every sampled rollout symmetrically.

By Chirag Chawla, Rohan Charudatt Salvi, Madhav S. Baidya