arXiv AI

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

arXiv:2605. 12969v3 Announce Type: replace-cross Abstract: Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks.

arXiv AI
Jul 22

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts

arXiv:2607. 18470v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled rollouts can be compared within a group using task-provided correctness signals.

By Yuxin Xiong, Xunyi Jiang, Rohan Surana, Xintong Li, Sheldon Yu, Nikki Lijing Kuang, Ryan A. Rossi, Jingbo Shang, Tong Yu, Julian McAuley, Junda Wu
arXiv AI
2d ago

Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States

The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.

By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
arXiv AI
Sep 15

HISPO: Hierarchical Importance-Sampling Policy Optimization with Entropy-Derived Segments

HISPO introduces a segment‑level policy‑optimization technique for reinforcement learning with verifiable rewards, creating entropy‑derived contiguous segments during rollout and applying clipped importance‑sampling correction at this granularity. It bridges the gap between token‑level and sequence‑level corrections, offering a middle‑ground approach for credit assignment in long‑form mathematical reasoning. Evaluated on Qwen3‑1.7B‑Base across six benchmarks, HISPO consistently outperforms or matches the strongest baselines in Pass@8 and Acc@8 metrics, notably improving AIME25 scores over GRPO and GSPO.

By Quoc-Vinh Lai-Dang, Hyo-Sang Shin
arXiv AI
Jul 7

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

arXiv:2607. 03126v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, Yanhua Cheng, Peng Jiang, Yadong Mu