Hugging Face Trending Papers

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Read the original on Hugging Face Trending Papers →

Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 1

Investigating Social Bias Changes in Quantized Language Models

Post‑training quantization of large language models reduces memory usage but can alter social biases in ways that aggregate metrics miss. In a large‑scale study of 50 quantized models on PostTrainingBiasBench, the authors discovered a phenomenon called quantization‑induced bias flipping, where up to 21% of responses switch from biased to unbiased or vice versa, especially for uncertain predictions and stronger quantization (4‑bit vs 8‑bit). These flips lead to asymmetric impacts across demographic groups, with some groups experiencing up to an 18.6% worsening of bias while others improve by 14.1%, resulting in misleadingly neutral overall scores.

By Stanley Z. Hua, Sanae Lotfi, Irene Y. Chen
arXiv Machine Learning
Sep 11

Why Does Post-Training Quantization Work?

Post‑training quantization compresses large language models by storing weights at reduced precision, introducing errors into hidden states that could accumulate with depth. However, pretrained models accumulate far less hidden‑state error than randomly initialized ones, largely preserving downstream performance. The study identifies two key mechanisms: (1) each layer’s new error tends to oppose inherited error, partially canceling it, and (2) the LM‑head geometry preserves high‑rank token scores, mitigating output changes.

By Yuxiang Chen, Michael Beyer, Jun Zhu, Jianfei Chen
arXiv Machine Learning
Sep 2

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

The paper investigates where post‑training quantization (PTQ) harms large language models (LLMs) and how to best allocate a limited precision budget. By causally raising each layer to 8‑bit precision across nine open‑weight models, the authors find that quantization damage is diffuse rather than concentrated in specific task circuits or weight statistics, and that globally refining quantization granularity outperforms selectively protecting the most recoverable layers. They also observe that the residual accuracy loss is budget‑limited and that peak recovery locations correlate with architecture within families but not across families.

By Jundong Hu, Shekar Ramachandran