Investigating Social Bias Changes in Quantized Language Models
Read the original on arXiv Computation and Language →Post‑training quantization of large language models reduces memory usage but can alter social biases in ways that aggregate metrics miss. In a large‑scale study of 50 quantized models on PostTrainingBiasBench, the authors discovered a phenomenon called quantization‑induced bias flipping, where up to 21% of responses switch from biased to unbiased or vice versa, especially for uncertain predictions and stronger quantization (4‑bit vs 8‑bit). These flips lead to asymmetric impacts across demographic groups, with some groups experiencing up to an 18.6% worsening of bias while others improve by 14.1%, resulting in misleadingly neutral overall scores.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.