arXiv AI

GRPODropout: Less is More for Online Reinforcement Learning Rollouts

arXiv Machine Learning
Jun 2

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

arXiv:2605. 21125v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models (LLMs).

By Xixiang He, Qiyao Sun, Ao Cheng, Xingming Li, Xuanyu Ji, Hailun Lu, Runke Huang, Qingyong Hu
arXiv Machine Learning
Jun 5

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models

arXiv:2606. 05434v1 Announce Type: new Abstract: Group Relative Policy Optimisation (GRPO) has emerged as an effective reinforcement-learning algorithm for aligning language models on reasoning tasks, but it treats every token position and every sampled rollout symmetrically.

By Chirag Chawla, Rohan Charudatt Salvi, Madhav S. Baidya
arXiv Machine Learning
2d ago

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

RollVerify is a lightweight reinforcement learning framework that addresses the trade‑off between efficiency and accuracy in long‑tail rollout settings. It introduces an off‑policy shift metric (OPS) to quantify deviation in partially generated trajectories and uses sequence‑level and token‑level verification to truncate invalid suffixes before training. Experiments on mathematical reasoning tasks show that RollVerify matches on‑policy accuracy while cutting training costs, with preliminary evidence from code‑generation tasks.

By Yongqiang Yao, Jinru Tan, Kaihuan Liang, Zixin Yin, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu
arXiv Machine Learning
Jun 25

ExTra: Exploratory Trajectory Optimization for Language Model Reinforcement Learning

arXiv:2606. 24994v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) for language-model reasoning can fail at both extremes of task difficulty: easy prompts often produce all-correct, low-diversity rollout groups with little gradient signal, while hard prompts can produce all-incorrect groups with no positive reward.

By Wenyang Hu, Junxiang Jia, Zhen Shu, Daniel Dahlmeier, See-Kiong Ng, Bryan Kian Hsiang Low
arXiv AI
5d ago

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT

The paper introduces Follow the Winners (FTW), a critic‑free policy‑learning algorithm that adapts the cross‑entropy method for reinforcement fine‑tuning of agentic large language models. FTW replaces group rollouts with an ordinal filter on replay‑buffer samples, achieving polynomial concentration in the order statistic of returns and offering a bounded risk‑seeking offset that balances variance reduction. Experiments on Sokoban and Search‑R1 show that FTW matches the performance of GRPO and PPO while reducing reliance on value models or repeated rollouts.

By Joery Ari\"en de Vries, Neil David Lawrence, Zhenwen Dai
arXiv AI
Oct 2

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo