arXiv Machine Learning By Sanae Lotfi, Polina Kirichenko, Steven Li, Zechun Liu

Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

Read the original on arXiv Machine Learning →

arXiv:2606. 00206v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
2d ago

RATIO: Reasoning Analysis and Token-level Inference Optimization for Quantized Reasoning Models

The paper introduces RATIO, a framework for improving quantized reasoning models by identifying overthinking tokens and applying token-specific penalties. It uses Quantization-aware Reasoning Behavior Analysis to detect problematic tokens and Token-Specific Penalty Determination to assign penalties without extra training. Experiments show RATIO outperforms existing methods, boosting accuracy by up to 9.8 points and shortening chain-of-thought length by up to 51.3%.

By Chengzhu Bao, Xianglong Yan, Tianao Zhang, Jiaqi Chen, Shaoqiu Zhang, Yulun Zhang
arXiv Machine Learning
Jun 25

Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models

arXiv:2606. 25519v1 Announce Type: cross Abstract: Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured by final-answer accuracy or per-token latency.

By Xinyu Lian, Walid Krichene, Beichen Huang, Masahiro Tanaka, Olatunji Ruwase, Li Zhang, Minjia Zhang