ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.
arXiv:2608. 07019v1 Announce Type: new Abstract: Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models.
arXiv:2608. 15567v1 Announce Type: new Abstract: Weight-only post-training quantization (PTQ) enables the deployment of large language models under tight memory budgets, but accuracy often collapses at 2-3 bits.
arXiv:2605.11222v2 Announce Type: replace Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
arXiv:2605. 08692v2 Announce Type: replace Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference.
arXiv:2607. 23047v1 Announce Type: cross Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget.
REAL-Q introduces a new post‑training quantization approach for large language models that replaces the traditional single closed‑form second‑order solver with a fine‑grained, dynamic block‑wise gradient descent applied after every 128‑column block. By aligning the surrogate loss with the end‑to‑end objective and using a sliding window for smooth cross‑layer transitions, REAL‑Q mitigates error propagation and information misalignment. Experiments on LLaMA‑3.1 and Qwen3 show up to ~49% reduction in end‑to‑end KL divergence compared to state‑of‑the‑art methods.
arXiv:2606. 10890v1 Announce Type: cross Abstract: Post-training quantization (PTQ) compresses large language models by mapping weights to low-bit representations.
arXiv:2602. 05367v3 Announce Type: replace Abstract: Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency and performance.
arXiv:2608.30564v1 Announce Type: cross Abstract: Mixed-precision quantization (MPQ) assigns a different bitwidth to each linear layer of a large language model (LLM) to minimize the quantization-ind...
arXiv:2606. 10531v1 Announce Type: cross Abstract: Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs).
arXiv:2602.03537v2 Announce Type: replace Abstract: Matryoshka Quantization (MatQuant), Any-Precision-LLM (AP) and AnyBCQ (AB) are recent quantization approaches showing that a single integer-quantiz...
arXiv:2609.00224v1 Announce Type: cross Abstract: Weight-only post-training quantization (PTQ) can alleviate the computational burden of serving large language models (LLMs) at scale. However, existi...