ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.
By Mehdi Makni, Ryan Lucas, Rahul Mazumder
Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment. Ternarization has emerged as a promising compression technique, offering significant reductions in model size and inference complexity.
arXiv:2606. 13054v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deployment.
By Zhixiong Zhao, Zukang Xu, Zhixuan Chen, Xing Hu, Zhe Jiang, Dawei Yang
arXiv:2605. 08692v2 Announce Type: replace Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference.
By Beshr IslamBouli, David Jin
arXiv:2507. 23035v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive capabilities across a wide range of applications, but demand substantial memory and compute resources during inference.
By Xueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu, Qilin Zheng, Yintao He, Hai Li
arXiv:2606. 07116v1 Announce Type: cross Abstract: Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost and memory usage.
By Haoqi Wang, Lorenz K. Mueller, Jiawei Zhuang, Mathieu Salzmann, Lukas Cavigelli