ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.
arXiv:2608. 15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits.
arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.
As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model...
arXiv:2608. 13966v1 Announce Type: new Abstract: As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality.
arXiv:2606. 10890v1 Announce Type: cross Abstract: Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations.
arXiv:2605.11222v2 Announce Type: replace Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
arXiv:2609.38599v1 Announce Type: new Abstract: Group-wise post-training quantizers for large language models round weights onto a grid that is not refit to the resulting integer codes. We show that...
G$^2$PTQ is a post‑training quantization framework that improves large language models by combining first‑ and second‑order information in a globally supervised, block‑wise optimization. It refreshes gradient and Hessian estimates before each Transformer block and uses a trust‑region scaling mechanism to stabilize gradient steps, preventing exploding weight updates. The method achieves better alignment with full‑precision models and outperforms state‑of‑the‑art baselines across various model families and bit‑widths.
arXiv:2607. 07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining.
arXiv:2601. 21626v2 Announce Type: replace-cross Abstract: Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing quantization error.
arXiv:2609.00224v1 Announce Type: cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existi...
arXiv:2609.06161v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved remarkable progress, yet their massive storage and memory-bandwidth demands still hinder efficient deploym...
Post-training quantization lowers the memory footprint of Large Language Models (LLMs) and speeds up inference, which is why it is now common for on-device deployment. Most of what we know about its e...