arXiv:2606. 13316v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs).
By Xucong Wang, Ziyu Ma, Yong Wang, Shidong Yang, Hailang Huang, Renda Li, Pengkun Wang, Xiangxiang Chu
arXiv:2606. 03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (CoTs).
By Ziyan Liu, Xueda Shen, Yuzhe Gu, Songyang Gao, Kuikun Liu, Guangran Cheng, Chengqi Lyu, Dahua Lin, Wenwei Zhang, Kai Chen
arXiv:2607. 28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring.
By Fei Ding, Yongkang Zhang, Runhao Liu, Yuhao Liao, Zijian Zeng
SPIRAL is a reinforcement‑learning framework that trains language models to employ three inference primitives—sequential reasoning within a trace, parallel sampling of independent traces, and aggregation of those traces—within a single compute pipeline. The model first generates multiple independent chain‑of‑thought traces in parallel, then produces a final aggregation trace conditioned on them, with all components optimized end‑to‑end for the reward of the aggregated response. Experiments on reasoning tasks demonstrate that SPIRAL scales efficiently with inference compute, achieving up to 11× better scaling efficiency and 15% higher performance compared to the GRPO baseline when all three primitives are scaled.
By Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman
arXiv:2606. 03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control.
By Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
arXiv:2609.37119v1 Announce Type: cross
Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
By Hongyang Li, Xiao Li, Caesar Wu, Said Mammar, Gr\'egoire Danoy, Pascal Bouvry
arXiv:2604.08454v2 Announce Type: replace
Abstract: Large language models are increasingly deployed in high-stakes domains, where confident yet incorrect inferences may cause severe real-world harm,...
By Haokai Ma, Lee Yan Zhen, Gang Yang, Yunxiang Chen, Yunshan Ma, Tat-Seng Chua, Ee-Chien Chang
Language model reasoning can be substantially improved at test time via scaffolds that scale inference compute across different primitives -- sequential reasoning within a trace, independently sampled parallel traces, and aggregation of multiple reasoning traces into a final response. During post-training, however, language models are optimized only for sequential reasoning within a single trace.
arXiv:2508. 02178v3 Announce Type: replace Abstract: Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and obscure the underlying reasoning process.
By Taihang Zhen, Jialiang Hong, Kai Chen, Guang Yang, Junlan Feng, Wenpeng Zhu, Jing Huo, Yang Gao, Depeng Wang, Haitao Wan, Xi Yang, Fanyu Meng, Yuyao Zhang, Ji Qi, Xiangyu Zhou
arXiv:2606. 11209v1 Announce Type: cross Abstract: Visual question answering increasingly requires multi-step reasoning.
By Jingpei Wu, Xiao Han, Weixiang Shen, Boer Zhang, Zifeng Ding, Volker Tresp
arXiv:2609.33149v2 Announce Type: replace
Abstract: A common principle of effective learning is to practice material that is neither already mastered nor too difficult to permit progress. We ask how...
By Hongbo Chen, Guohua Lu, Ting Dang, Hong Jia
arXiv:2509. 04027v4 Announce Type: replace Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs).
By Zeyu Gan, Hao Yi, Yong Liu