Hugging Face Trending Papers

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization.

arXiv Computation and Language
Sep 1

Investigating Social Bias Changes in Quantized Language Models

Post‑training quantization of large language models reduces memory usage but can alter social biases in ways that aggregate metrics miss. In a large‑scale study of 50 quantized models on PostTrainingBiasBench, the authors discovered a phenomenon called quantization‑induced bias flipping, where up to 21% of responses switch from biased to unbiased or vice versa, especially for uncertain predictions and stronger quantization (4‑bit vs 8‑bit). These flips lead to asymmetric impacts across demographic groups, with some groups experiencing up to an 18.6% worsening of bias while others improve by 14.1%, resulting in misleadingly neutral overall scores.

By Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen
arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
arXiv Machine Learning
Sep 2

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

The paper investigates where post‑training quantization (PTQ) harms large language models (LLMs) and how to best allocate a limited precision budget. By causally raising each layer to 8‑bit precision across nine open‑weight models, the authors find that quantization damage is diffuse rather than concentrated in specific task circuits or weight statistics, and that globally refining quantization granularity outperforms selectively protecting the most recoverable layers. They also observe that the residual accuracy loss is budget‑limited and that peak recovery locations correlate with architecture within families but not across families.

By Jundong Hu, Shekar Ramachandran
arXiv AI
Sep 10

Accuracy is Not Enough: A Divergence-Based Approach to Evaluate Fidelity Loss in Quantized LLMs

The paper argues that relying solely on zero‑shot task accuracy is insufficient for evaluating quantized large language models (LLMs) because accuracy ignores changes in the full predictive distribution. It proposes a distribution‑sensitive framework that measures fidelity loss by computing statistical distances—such as Jensen‑Shannon Divergence and Total Variation Distance—between the full‑vocabulary output distributions of a full‑precision BF16 reference and its quantized counterparts. Experiments across five foundation architectures and four reasoning benchmarks show that these divergence metrics increase with stronger quantization, revealing distributional drift that top‑1 accuracy fails to capture, and suggest that mixed‑precision Q4_K schemes can offer lower divergence than uniform Q4_0 at comparable memory usage.

By Shahzeb Qamar, Lorenz Sparrenberg, Christian Bauckhage, Baha Rababah, Carson Leung, Murat Kantarcioglu, Cuneyt Gurcan Akcora, Rafet Sifa
arXiv Machine Learning
Aug 20

Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs

The study investigates how post‑training quantization (PTQ) affects proactive interference (PI) in large language models. Using bitsandbytes, the authors compare FP16, INT8, and INT4/NF4 precision across three instruction‑tuned models and find that INT4 quantization markedly degrades accuracy under high interference, with INT8 also incurring a smaller penalty in two of the three models. The degradation is linked to increased same‑key intrusion errors and originates in the quantized transformer backbone rather than the output layer.

By Shayan Shahrabi-Farahani (Shahid Beheshti University, Tehran, Iran), Dara Rahmati (Shahid Beheshti University, Tehran, Iran)
arXiv AI
Aug 26

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

The paper investigates how quantization affects large language models’ self‑explanations, examining natural language explanations and counterfactual examples across three quantization techniques and bit widths. Results show moderate declines in explanation quality (up to 4.4%) and faithfulness (up to 3.9%), with user studies indicating up to an 8.5% drop in coherence and trustworthiness. Larger models are less resilient in quality but remain more faithful, and no single quantization method consistently outperforms others across accuracy, quality, and faithfulness.

By Qianli Wang, Nils Feldhus, Pepa Atanasova, Fedor Splitt, Simon Ostermann, Sebastian M\"oller, Vera Schmitt