arXiv:2608. 03068v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs).
By Ziqi Jia, Yalu Ouyang, Bo Pang, Panpan Li, Hangfei Xu, Shengzhao Wen, Shiyong Li, Yanpeng Wang
Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). However, existing methods suffer from insufficient precision in feedback on generated answer trajectories and exhibit the phenomenon of problem difficulty drift.
arXiv:2606. 25178v2 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science.
By Yongjin Yang, Jiarui Liu, Yinghui He, Lechen Zhang, Bernhard Sch\"olkopf, Zhijing Jin
arXiv:2509. 25004v2 Announce Type: replace Abstract: Online reinforcement learning with verifiable rewards (RLVR) has become an effective paradigm for improving the reasoning abilities of large language models, but most methods still optimize reasoning trajectories over the static problem set, wasting rollout budget on solved or overly difficult problems.
By Shijie Zhang, Zheng Xiao, Shiyu Liu, Guohao Sun, Kevin Zhang, Xiang Guo, Rujun Guo, Shaoyu Liu, Wangxiao Zhao, Guanjun Jiang
TTSR (Test-Time Self-Reflection) is a framework that enables large language models to adapt during inference by alternating between a Student role that solves test questions and a Teacher role that analyzes failures and generates targeted variant questions. The method incorporates a weakness memory and a strategy note to guide exploration, reducing reliance on noisy pseudo-labels and inefficient rollouts. Experiments on mathematical reasoning benchmarks demonstrate consistent test-time improvements, strong cross-backbone generalization, and transfer to general-domain reasoning tasks.
By Haoyang He, Zihua Rong, Yunjia Zhao, Lan Yang, Jian Chang, Honggang Zhang
arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.
By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation proposes a plug‑and‑play graph‑based online difficulty estimator for reinforcement learning with verifiable rewards (RLVR). The method constructs a difficulty‑aware sample graph using semantic and reasoning similarities, introduces latent difficulty states with a Potts prior, aggregates rollout outcomes with a state‑level Beta‑Binomial model, and updates these estimates online via a mean‑field variational algorithm. This framework can be integrated into sample‑selection and rollout‑allocation schedulers, enabling difficulty‑adaptive exploration without dedicated probing and achieving better performance across multiple base models, RL schedulers, and benchmarks.
By Zhizhao Liu, Zhiliang Tian, Xi Wang, Zhihua Wen, Yihang Xiong, Zhiquan Lai, Dongsheng Li
arXiv:2609.13997v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Ye...
By Yukang Zhu, Zhen Han
On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting.
Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization. Existing task-valuation methods mostly rely on snapshot-based signals such as current pass rate or reward, which estimate how solvable a task is under the current policy.
arXiv:2608.16647v2 Announce Type: replace
Abstract: On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalizati...
By Zhaoyi Li, Deyang Kong, Yuan Wei, Evan Yang, Ranran Shen, Mahardika Krisna Ihsani, Ming Yang, Wei Zhang, Chuan Hao, Jian Yang, Ran Tao, Bryan Dai, Shikun Zhang, Wei Ye, Ying Wei, Defu Lian
arXiv:2606. 19750v1 Announce Type: cross Abstract: Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training efficiency depends critically on how problems are sampled during optimization.
By Darrien McKenzie, Nicklas Hansen, Xiaolong Wang