arXiv:2606. 26997v1 Announce Type: cross Abstract: Large language model (LLM) post-training for reasoning increasingly relies on reinforcement learning with verifiable rewards (RLVR), where models learn from ground-truth feedback on mathematical, logical, and scientific tasks.
By Rongjian Chen, Jianmin Hu, Kejiang Ye, Minxian Xu
The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.
By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
arXiv:2607. 20438v1 Announce Type: cross Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque.
By Peiyan Zhang, Haibo Jin, Liying Kang, Haohan Wang
arXiv:2606. 18521v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Reward (RLVR) has emerged as a powerful post-training paradigm that surpasses Supervised Fine-Tuning (SFT) in eliciting reasoning intelligence and resisting catastrophic forgetting.
By Chenrui Wu, Zexi Li, Jiajun Bu, Jiangchuan Liu, Haishuai Wang
arXiv:2608.24696v1 Announce Type: cross
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training lar...
By Wenze Lin, Jiale Zhao, Xitai Jiang, Songde Rao, Yining Li, Shenzhi Wang, Bingxiang He, Gao Huang
arXiv:2607. 01232v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers.
By Zijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong