arXiv Machine Learning By Xinyu Lian, Walid Krichene, Beichen Huang, Masahiro Tanaka, Olatunji Ruwase, Li Zhang, Minjia Zhang

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

Read the original on arXiv Machine Learning →

arXiv:2606. 25519v1 Announce Type: cross Abstract: Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
2d ago

RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

The paper introduces RATIO, a framework for improving quantized reasoning models by identifying overthinking tokens and applying token-specific penalties. It uses Quantization-aware Reasoning Behavior Analysis to detect problematic tokens and Token-Specific Penalty Determination to assign penalties without extra training. Experiments show RATIO outperforms existing methods, boosting accuracy by up to 9.8 points and shortening chain-of-thought length by up to 51.3%.

By Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang