Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates post‑training quantization of transformer attention blocks by optimizing a joint loss over the Q, K, V projections rather than individual weight matrices. Using this joint attention‑based objective (JAB), the authors achieve significant compression on Mistral‑7B, recovering 77‑90% of the performance gap at 3 bits, but the method fails when MLP layers are included. A role‑aware offset rule that ignores sensitivity estimates outperforms JAB on GPT‑2 and full Mistral‑7B, demonstrating that the matrix a weight belongs to is more critical than sensitivity metrics.
The paper investigates where post‑training quantization (PTQ) harms large language models (LLMs) and how to best allocate a limited precision budget. By causally raising each layer to 8‑bit precision across nine open‑weight models, the authors find that quantization damage is diffuse rather than concentrated in specific task circuits or weight statistics, and that globally refining quantization granularity outperforms selectively protecting the most recoverable layers. They also observe that the residual accuracy loss is budget‑limited and that peak recovery locations correlate with architecture within families but not across families.
arXiv:2606. 04620v1 Announce Type: cross Abstract: LLMs have become the state-of-the-art algorithms for solving NLP tasks.
arXiv:2607. 10137v1 Announce Type: new Abstract: Post-training quantization (PTQ) of large language models degrades sharply below 4-bit precision.
arXiv:2608. 08188v1 Announce Type: new Abstract: Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determined by bit-width alone.
arXiv:2607. 00908v1 Announce Type: new Abstract: Mixed-precision quantization (MPQ) has become a key technique for deploying large language models under stringent memory and compute constraints.