arXiv Machine Learning By Fei Ding, Yongkang Zhang, Yuhao Liao, Zijian Zeng, Huiming Yang

On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

Read the original on arXiv Machine Learning →

arXiv:2607. 23364v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.