← Back to all news
arXiv Machine Learning August 10, 2026 By Zekun Wu, Swati Dhiman, Adriano Koshiyama

Quantization Damage Is Multiplicative, Not Additive

Read the original on arXiv Machine Learning →

arXiv:2608. 06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

  • llms
  • agents
  • efficiency
  • benchmarks
  • safety

Related stories

arXiv Machine Learning
Aug 11

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

arXiv:2608. 06564v2 Announce Type: replace Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts.

By Zekun Wu, Swati Dhiman, Adriano Koshiyama
llmsagentsefficiencybenchmarkssafety
More like this →
arXiv Machine Learning
Jul 31

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

arXiv:2607. 27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless.

By Jiwon Jang, Kisu Yang, Heuiseok Lim, Hyunwoo Park
llmsagentsefficiencybenchmarkssafety
More like this →
arXiv Machine Learning
Jul 21

Half the Experts, All the Code: One-Shot Domain Pruning of Mixture-of-Experts LLMs for Coding

arXiv:2607. 16721v1 Announce Type: new Abstract: The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of "expert" subnetworks, of which only a few act on any token.

By Anik Jha
llmsagentsfine-tuningefficiencybenchmarks
More like this →
arXiv AI
Jun 3

How Quantization Changes Interpretable Features: A Sparse Autoencoder Analysis of Language Models

arXiv:2606. 03002v1 Announce Type: cross Abstract: Quantization is a standard path to deploying large language models, and a quantized model is typically judged acceptable when its perplexity or downstream accuracy stays close to the full-precision original.

By Evan Duan
llmsefficiencysafety
More like this →
arXiv Machine Learning
Jul 15

Saturation Makes Quantization Error Additive: A Coverage Model with a Certificate

arXiv:2607. 12266v1 Announce Type: new Abstract: Mixed-precision quantization must decide which parts of a model to keep at higher precision.

By Joshua Hill
efficiency
More like this →
arXiv AI
Jun 8

Perplexity Can Miss SAE Feature Damage Under Quantization

arXiv:2606. 03002v2 Announce Type: replace-cross Abstract: Quantization is a standard path to deploying large language models, and quantized models are typically judged acceptable when perplexity or downstream accuracy remains close to the full-precision original.

By Evan Duan
llmsefficiencysafety
More like this →