arXiv Machine Learning By Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals

Read the original on arXiv Machine Learning →

Range-GRPO introduces a semi‑supervised post‑training framework that uses conformally calibrated reward ranges instead of single point scores for large language models. By comparing reward ranges pairwise within rollout groups, the method incorporates reward uncertainty into both the magnitude and direction of learning signals. Experiments show that Range‑GRPO outperforms other semi‑supervised approaches on both in‑distribution and out‑of‑distribution tasks while using fewer training resources.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 30

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

arXiv:2606. 28707v1 Announce Type: new Abstract: Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces memory and compute overhead relative to critic-based PPO pipelines for aligning large language models.

By Yupeng Chang, Yuan Wu, Yi Chang
arXiv AI
2d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo