arXiv:2509. 22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available.
By Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura, Mitsuki Sakamoto, Ryota Mitsuhashi, Eiji Uchibe
arXiv:2609.00213v1 Announce Type: new
Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific e...
By Yu Yuan, Yaoyou Fan, Lili Zhao, Guangting Zheng, Kai Zhang, Lu Pan, Ke Zeng, Qi Liu
The paper introduces Density-Aware Reward Aggregation (DARA), a method that adjusts reward weighting in multi-reward reinforcement learning based on the density of active rewards within rollout batches. By applying an inverse-square-root density correction, DARA gives more weight to less frequently active rewards, enabling faster learning of targeted behaviors. Experiments on tool calling and mathematical reasoning demonstrate that DARA achieves comparable final performance while reducing training steps by up to 26% and 65% respectively.
By Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu
The paper investigates why Group Relative Policy Optimization (GRPO) benefits from per‑prompt normalization by examining the local curvature of the sequence‑level policy gradient. It shows that standard deviation normalization acts as an adaptive gradient, yielding a provably faster convergence rate than unnormalized REINFORCE under mild conditions, with the improvement tied to the average within‑prompt reward standard deviation. The authors also propose IS‑GRPO, an importance‑sampling variant that maintains alignment with the full gradient and offers a tighter convergence guarantee, and empirically validate these theoretical insights on GSM8K and MATH datasets at 1.5B and 7B model scales.
By Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang
arXiv:2606. 26300v1 Announce Type: new Abstract: A classical intuition holds that verifying a solution is easier than producing one.
By Binghai Wang, Chenlong Zhang, Dayiheng Liu, Jiajun Zhang, Jiawei Chen, Mouxiang Chen, Rongyao Fang, Siyuan Zhang, Xuwu Wang, Yuheng Jing, Zeyao Ma, Zeyu Cui
Range-GRPO introduces a semi‑supervised post‑training framework that uses conformally calibrated reward ranges instead of single point scores for large language models. By comparing reward ranges pairwise within rollout groups, the method incorporates reward uncertainty into both the magnitude and direction of learning signals. Experiments show that Range‑GRPO outperforms other semi‑supervised approaches on both in‑distribution and out‑of‑distribution tasks while using fewer training resources.
By Ryunyi Lee, Kangjun Noh, Somin Kim, Heedong Kim, Kyungwoo Song
arXiv:2507. 01551v3 Announce Type: replace Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs).
By Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong
arXiv:2602.05863v3 Announce Type: replace
Abstract: Group Relative Policy Optimization (GRPO) remains the dominant critic-free approach for fine-tuning LLMs and VLMs, but its compatibility with const...
By Roger Girgis, Rodrigue de Schaetzen, Luke Rowe, Azal\'ee Robitaille, Christopher Pal, Liam Paull
The paper introduces Gradient-Aligned Reward (GAR), a reinforcement learning technique that uses truncated backpropagation to generate a compact gradient vector for each rollout and compares it to an expert-anchor gradient via cosine similarity. This dense, reasoning-aware reward improves large language model chain-of-thought reasoning on math benchmarks and transfers to other tasks without domain‑specific data, while adding less than 9% computational overhead. GAR outperforms existing baselines such as GRPO on Qwen3-4B and Qwen3-8B models.
By Leqi Zheng, Jinbo Su, Fang Niu, Chaokun Wang, Weiping Wang, Jiajun Zhang, Shannan Yan, Jie Wu, Zhaolu Kang, Rong Fu, Hang Zhang
arXiv:2606. 15385v1 Announce Type: new Abstract: Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a central challenge in AI safety.
By \"Omer Veysel \c{C}a\u{g}atan, Xuandong Zhao
arXiv:2606. 16771v1 Announce Type: new Abstract: As LLMs advance, post-training reinforcement learning (RL) increasingly relies on multi-dimensional rewards to cultivate comprehensive capabilities.
By Haotian Liu, Yihao Liu, Jingwei Ni, Siyuan Huang, Xinpeng Liu, Pengyu Cheng, Jiajun Song, Ruijin Ding, Junfeng Li, Zhechao Yu, Mengyu Zhou, Hongteng Xu, Xiaoxi Jiang, Guanjun Jiang
Group Relative Policy Optimization (GRPO) assigns a magnitude to each rollout based on within‑group reward statistics, rewarding rollouts that reach the correct answer through reasoning. However, the same magnitude can be high for rollouts that reach the answer by guessing, creating a spurious advantage that misleads the policy toward guess‑like behaviors. The paper identifies three scenarios where this occurs—bounded‑answer tasks, open‑answer sets with bounded sub‑cases, and search agents with many paths to the same answer—and proposes SIGNBALANCE, a composition‑free magnitude that preserves the verifier sign, uses a global scale, and restores zero‑mean balance via stop‑gradient per‑class rescaling, matching GRPO on open‑answer math and improving on bounded‑answer math and search agents.
By Jiamian Wang, Samyadeep Basu, Koustava Goswami, Tong Yu, Zhiqiang Tao