← Back to all news
arXiv Computation and Language September 22, 2026 By Peidong Wang, Demi Wang, Xufang Luo, Jiahang Xu, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li

What are Key Factors for Updates in RL for LLM Reasoning?

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • reinforcement-learning
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 2

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning

arXiv:2606. 01281v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs).

By Yixiu Mao, Yun Qu, Qi Wang, Heming Zou, Xiangyang Ji
llmsreinforcement-learningfine-tuningbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 30

Experience Augmented Policy Optimization for LLM Reasoning

arXiv:2606. 30420v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs).

By Jinda Lu, Kexin Huang, Junkang Wu, Shuo Yang, Jinghan Li, Chiyu Ma, Shaohang Wei, Xiang Wang, Guoyin Wang, Jingren Zhou
llmsreinforcement-learningbenchmarks
More like this →
arXiv AI
Jul 1

Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index

arXiv:2606. 31575v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards more robust reasoning capabilities.

By Outongyi Lv, Yanzhao Zheng, Yuanwei Zhang, Zhenghao Huang, Xingjun Wang, Baohua Dong, Hangcheng Zhu, Yingda Chen
llmsreinforcement-learningbenchmarks
More like this →
arXiv Machine Learning
Jun 4

Policy Improvement Reinforcement Learning

arXiv:2604. 00860v3 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a central post-training paradigm for improving the reasoning capabilities of large language models.

By Huaiyang Wang, Xiaojie Li, Deqing Wang, Haoyi Zhou, Zixuan Huang, Yaodong Yang, Jianxin Li, Yikun Ban
llmsreinforcement-learningbenchmarks
More like this →
arXiv AI
Aug 11

An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning

arXiv:2504. 18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated by systems such as OpenAI's O1~\cite{o1} and DeepSeek-R1~\cite{r1}.

By Tianbing Xu
llmsreinforcement-learningsafety
More like this →
arXiv Machine Learning
Jul 9

UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma

arXiv:2607. 06987v1 Announce Type: new Abstract: Reinforcement learning (RL) has become the standard paradigm for enhancing the complex reasoning capabilities of large language models (LLMs).

By Chongyu Fan, Pengfei Liu, Jingjia Huang, Sijia Liu, Yi Lin
llmsreinforcement-learningmultimodal
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea