ReCo: Reweighting GRPO Against Distributional Concentration
arXiv:2607. 26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models.
arXiv:2607. 26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models.
arXiv:2605. 21125v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models (LLMs).
arXiv:2603. 25184v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks.
arXiv:2606. 05434v1 Announce Type: new Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks, but it treats every token position and every sampled rollout symmetrically.
RollVerify is a lightweight reinforcement learning framework that addresses the trade‑off between efficiency and accuracy in long‑tail rollout settings. It introduces an off‑policy shift metric (OPS) to quantify deviation in partially generated trajectories and uses sequence‑level and token‑level verification to truncate invalid suffixes before training. Experiments on mathematical reasoning tasks show that RollVerify matches on‑policy accuracy while cutting training costs, with preliminary evidence from code‑generation tasks.
arXiv:2606. 24994v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward.
arXiv:2607. 16205v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which typically optimizes the policy by contrasting multiple self generated rollouts.
The paper introduces Follow the Winners (FTW), a critic‑free policy‑learning algorithm that adapts the cross‑entropy method for reinforcement fine‑tuning of agentic large language models. FTW replaces group rollouts with an ordinal filter on replay‑buffer samples, achieving polynomial concentration in the order statistic of returns and offering a bounded risk‑seeking offset that balances variance reduction. Experiments on Sokoban and Search‑R1 show that FTW matches the performance of GRPO and PPO while reducing reliance on value models or repeated rollouts.
arXiv:2605.27293v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existin...
arXiv:2608. 01717v1 Announce Type: new Abstract: Recent reinforcement learning methods for diffusion large language models (dLLMs) commonly rely on on-policy rollouts generated by the target dLLM itself.
arXiv:2607. 24833v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from severe reward sparsity on competition-level mathematics.
The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.