arXiv AI By Baha Rababah, Cuneyt Gurcan Akcora, Carson K. Leung

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Read the original on arXiv AI →

arXiv:2607. 08734v1 Announce Type: new Abstract: Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Investigating Social Bias Changes in Quantized Language Models

Post‑training quantization of large language models reduces memory usage but can alter social biases in ways that aggregate metrics miss. In a large‑scale study of 50 quantized models on PostTrainingBiasBench, the authors discovered a phenomenon called quantization‑induced bias flipping, where up to 21% of responses switch from biased to unbiased or vice versa, especially for uncertain predictions and stronger quantization (4‑bit vs 8‑bit). These flips lead to asymmetric impacts across demographic groups, with some groups experiencing up to an 18.6% worsening of bias while others improve by 14.1%, resulting in misleadingly neutral overall scores.

By Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen
arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
arXiv AI
Sep 10

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

The paper argues that relying solely on zero‑shot task accuracy is insufficient for evaluating quantized large language models (LLMs) because accuracy ignores changes in the full predictive distribution. It proposes a distribution‑sensitive framework that measures fidelity loss by computing statistical distances—such as Jensen‑Shannon Divergence and Total Variation Distance—between the full‑vocabulary output distributions of a full‑precision BF16 reference and its quantized counterparts. Experiments across five foundation architectures and four reasoning benchmarks show that these divergence metrics increase with stronger quantization, revealing distributional drift that top‑1 accuracy fails to capture, and suggest that mixed‑precision Q4_K schemes can offer lower divergence than uniform Q4_0 at comparable memory usage.

By Shahzeb Qamar, Lorenz Sparrenberg, Christian Bauckhage, Baha Rababah, Carson Leung, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Rafet Sifa