arXiv:2606. 13316v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs).
By Xucong Wang, Ziyu Ma, Yong Wang, Shidong Yang, Hailang Huang, Renda Li, Pengkun Wang, Xiangxiang Chu
arXiv:2606. 03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs).
By Ziyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao, Kuikun Liu, Guangran Cheng, Chengqi Lyu, Dahua Lin, Wenwei Zhang, Kai Chen
arXiv:2607. 28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring.
By Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng
SPIRAL is a reinforcement‑learning framework that trains language models to employ three inference primitives—sequential reasoning within a trace, parallel sampling of independent traces, and aggregation of those traces—within a single compute pipeline. The model first generates multiple independent chain‑of‑thought traces in parallel, then produces a final aggregation trace conditioned on them, with all components optimized end‑to‑end for the reward of the aggregated response. Experiments on reasoning tasks demonstrate that SPIRAL scales efficiently with inference compute, achieving up to 11× better scaling efficiency and 15% higher performance compared to the GRPO baseline when all three primitives are scaled.
By Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman
arXiv:2606. 03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control.
By Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
arXiv:2609.37119v1 Announce Type: cross
Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
By Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry