arXiv:2606. 00494v1 Announce Type: new Abstract: Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment.
By Wneya Yu, Chao Zhang, Li Wang, Samson Lasaulce, Merouane Debbah
As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model...
arXiv:2608. 13966v1 Announce Type: new Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality.
By Vincent Counathe, Ben Athiwaratkun, Christopher De Sa, Tianyi Zhang
arXiv:2510. 18784v3 Announce Type: replace Abstract: Despite significant work on low-bit quantization-aware training (QAT), there is still an accuracy gap between such techniques and native training.
By Soroush Tabesh, Mher Safaryan, Andrei Panferov, Alexandra Volkova, Dan Alistarh
G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.
By Ruikang Liu, Haoli Bai, Yuxuan Sun, Qian Zhang, Wenzheng Cai, Yanqi Hao, Feiyu Wang, Weidong Zhong, Zhuang Wang, Tong Yang, Xiangsheng Zhou
arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.
By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv:2605.11222v2 Announce Type: replace
Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
By Ryan Lucas, Mehdi Makni, Xiang Meng, Adam Deng, Rahul Mazumder
arXiv:2606. 00573v1 Announce Type: new Abstract: Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices.
By Haiyu Wang, Yutong Wang, Leshu Li, Yihui Ren, Sai Qian Zhang
arXiv:2608.30384v1 Announce Type: new
Abstract: By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension,...
By Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato, Jiri Isa, Igor Ginzburg
arXiv:2601. 21626v2 Announce Type: replace-cross Abstract: Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing quantization error.
By Jinhao Zhang, Yunquan Zhang, Zicheng yan, Boyang Zhang, Jun Sun, Daning Cheng
arXiv:2607. 07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining.
By Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda
arXiv:2608. 14149v1 Announce Type: new Abstract: Recent training-free post-training quantization methods restore model accuracy through closed-form residual compensation.
By Lin-Fa Lee, Yi-Yu Chang, Kuo-Hei Yeh