The paper introduces POISE, a reinforcement learning algorithm that uses a model’s internal states as a value estimator to reduce variance in reinforcement learning with verifiable rewards (RLVR). By employing a lightweight probe that reads internal signals during the forward pass, POISE predicts baselines online and uses a cross‑rollout construction to keep gradients unbiased. Experiments on Qwen3‑4B and OLMo3‑7B‑Instruct‑DPO across six domains show POISE outperforms existing RLVR baselines, offering more stable training and a value model that generalizes across tasks and scales with the policy.
By Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo
arXiv:2608. 09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time.
By Vatsal Venkatkrishna, Nico Daheim, Iryna Gurevych
arXiv:2606. 01281v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs).
By Yixiu Mao, Yun Qu, Qi Wang, Heming Zou, Xiangyang Ji
arXiv:2603. 25184v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks.
By Jiahao Wu, Ning Lu, Shengcai Liu, Kun Wang, Yanting Yang, Bailong Lin, Chen Jason Zhang, Li Qing, Ke Tang
arXiv:2605.27293v2 Announce Type: replace
Abstract: Reinforcement learning with verifiable rewards has become a standard recipe for improving the reasoning abilities of large language models. Existin...
By Shijin Gong, Erhan Xu, Kai Ye, Giulia Livieri, Francesco Quinzan, Chengchun Shi
The paper introduces FastRL, a reinforcement learning framework designed to enhance the efficiency of Group Relative Policy Optimization (GRPO) and its variants. FastRL employs an advantage-aware pruning strategy that retains high-advantage trajectories while preserving gradient diversity, and an adaptive rollout sampling mechanism that adjusts sampling scale during training based on historical pruning data. Experiments show that FastRL can be integrated into GRPO, DAPO, and GSPO, yielding a 2.07× speedup on Geometry3K and GeoQA8K-R1V and a 1.64% accuracy improvement on visual reasoning benchmarks.
By Jiahua Yang, Zhiwei Yang, Xianpeng Zhang, Dongyu Chen, Xing Chen, Tianhuang Su, Haonan Lu, Quanlong Guan, Kai Tang, Chuangchuang Wang