arXiv:2608. 20256v1 Announce Type: new Abstract: Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones.
By Gijs Kassenaar, Zhao Yang, Vincent Fran\c{c}ois-Lavet
The paper introduces Budget‑Efficient Thinking (BET), a two‑stage framework that treats adaptive reasoning as a computational investment, aligning solve‑or‑fold decisions with expected return rather than perceived difficulty. BET learns three distinct behaviors: concise short solves for easy queries, early abstention (nice fold) when further reasoning is unlikely to pay off, and allocating sufficient compute (hero call) for hard‑but‑solvable questions. Experiments on seven benchmarks with three base models show BET cuts reasoning tokens by 54% while boosting accuracy by up to 3.2%, and it transfers effectively to scientific QA and logical reasoning tasks.
By Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu, Tingzhao Li, Yiqing Hu, Yumeng Zhao
The paper introduces CARE, a contrastive accuracy reward estimation method that adaptively adjusts reasoning length for large language models. By comparing beneficial length adjustments from online sampled responses, CARE applies adaptive length rewards within Group Relative Policy Optimization without extra hyperparameters or inference cost. Experiments on multiple reasoning benchmarks show that CARE improves Pass@1 by up to 4% while reducing reasoning length by 37%, achieving higher token efficiency.
By Zhengdong He, Yunfan Zhou, Jianguo Yao, Haibing Guan, Xijun Li
arXiv:2609.37304v1 Announce Type: new
Abstract: Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not...
By Zhibin Wen, Tao Han, Lei Bai, Can Li, Yang Xu
arXiv:2606. 07108v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking".
By Tengyao Tu, Yulin Li, Hui-Ling Zhen, Libo Qin, Zhoujun Wei, Jinghua Piao, Zhuotao Tian, Yong Li, Min Zhang
arXiv:2606. 07950v1 Announce Type: new Abstract: RL with verifiable rewards can substantially improve LLM reasoning, yet standard GRPO-style training often treats easy, hard, and learnable questions alike through uniform sampling and weighting, leading to inefficient compute allocation.
By Zhanke Zhou, Xiangyu Lu, Chentao Cao, Brando Miranda, Tongliang Liu, Bo Han, Sanmi Koyejo