Large Reasoning Models (LRMs) rely on long reasoning traces, making inference expensive. While low-bit quantization reduces per-token decoding cost, we show that aggressive 2-bit inference can fail to deliver end-to-end speedup because instability in the generation process inflates total token count.
arXiv:2609.26708v1 Announce Type: new
Abstract: Quantization-aware distillation (QAD) restores much of the short-form question-answering performance lost to sub-3-bit quantization, yet leaves mathema...
By Yuanteng Chen, Zhilei Liu, Peisong Wang, Yuantian Shao, Chuangyi Li, Weining Wang, Shuang Qiu, Gang Li, Jing Liu, Jian Cheng
arXiv:2606. 15682v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong problem-solving through long chain-of-thought, but their deployment is constrained by the high cost of full-precision inference and growing KV cache footprints.
By Janghwan Lee, Sihwa Lee, Jinseok Kim, Yongjik Kim, Jieun Lim, Jinwook Oh, Jungwook Choi
arXiv:2606. 25519v1 Announce Type: cross Abstract: Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency.
By Xinyu Lian, Walid Krichene, Beichen Huang, Masahiro Tanaka, Olatunji Ruwase, Li Zhang, Minjia Zhang
The paper introduces RATIO, a framework for improving quantized reasoning models by identifying overthinking tokens and applying token-specific penalties. It uses Quantization-aware Reasoning Behavior Analysis to detect problematic tokens and Token-Specific Penalty Determination to assign penalties without extra training. Experiments show RATIO outperforms existing methods, boosting accuracy by up to 9.8 points and shortening chain-of-thought length by up to 51.3%.
By Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang
TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.
By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher
arXiv:2606. 00206v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood.
By Sanae Lotfi, Polina Kirichenko, Steven Li, Zechun Liu
arXiv:2606. 26488v1 Announce Type: new Abstract: Recursive reasoning models can solve complex structured tasks with only a few million parameters by repeatedly updating a latent state.
By Pearse Jim, Steven Kolawole, Opegbemi Matthias Busoye, Glory Bagai, Virginia Smith
arXiv:2609.37304v1 Announce Type: new
Abstract: Large reasoning models improve performance on challenging problems by allocating additional computation before answering, but longer reasoning does not...
By Zhibin Wen, Tao Han, Lei Bai, Can Li, Yang Xu
arXiv:2608. 15797v1 Announce Type: new Abstract: KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history.
By Minsoo Cheong, Woosang Lim, Vincent-Daniel Yun, Sungjoo Yoo
arXiv:2606. 13233v1 Announce Type: cross Abstract: Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs.
By Sihwa Lee, Janghwan Lee, Donghoon Yoo, Jae Gon Kim, Hanyul Ryu, Soojung Ryu, Jungwook Choi
arXiv:2605.28006v2 Announce Type: replace-cross
Abstract: Understanding how LLMs reason is hindered by a practical asymmetry: while their generated outputs are observable, the underlying reasoning pa...
By Leonardo Matthew Yauw, Wei-Bin Kou, Yujiu Yang