← Back to all news
arXiv Machine Learning August 11, 2026 By Zekun Wu, Swati Dhiman, Adriano Koshiyama

Which Decisions Low-Bit Quantization Breaks, and How to Predict Them

Read the original on arXiv Machine Learning →

arXiv:2608. 06564v2 Announce Type: replace Abstract: Quantization is how large language models are actually deployed, and below four bits it hurts.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

  • llms
  • agents
  • efficiency
  • benchmarks
  • safety

Related stories

arXiv Machine Learning
Aug 10

Quantization Damage Is Multiplicative, Not Additive

arXiv:2608. 06564v1 Announce Type: new Abstract: Quantization is how large language models are actually deployed, and below four bits it is known to hurt.

By Zekun Wu, Swati Dhiman, Adriano Koshiyama
llmsagentsefficiencybenchmarkssafety
More like this →
arXiv Machine Learning
Jul 21

Half the Experts, All the Code: One-Shot Domain Pruning of Mixture-of-Experts LLMs for Coding

arXiv:2607. 16721v1 Announce Type: new Abstract: The strongest open-weight coding models are mixture-of-experts (MoE) networks: most of their size comes from large pools of "expert" subnetworks, of which only a few act on any token.

By Anik Jha
llmsagentsfine-tuningefficiencybenchmarks
More like this →
arXiv Machine Learning
Jul 31

Flat Score, Amplified Failures: How the Error Budget Masks Damage in Quantized LLM Agents

arXiv:2607. 27275v1 Announce Type: new Abstract: Post-training quantization to 4-bit weights is widely reported to be nearly lossless.

By Jiwon Jang, Kisu Yang, Heuiseok Lim, Hyunwoo Park
llmsagentsefficiencybenchmarkssafety
More like this →
arXiv AI
Aug 3

When Bits Break Recourse: Counterfactual-Faithful Quantization

arXiv:2605. 17160v2 Announce Type: replace-cross Abstract: Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is preserved.

By Chaymae Yahyati, Ismail Lamaakal, Khalid El Makkaoui, Ibrahim Ouahbi
efficiency
More like this →
arXiv AI
Jun 30

The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis

arXiv:2606. 29581v1 Announce Type: cross Abstract: Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluations usually treat these choices as fixed implementation details.

By Hari Prasad, Ritam Pal
llmsefficiencymultimodalsafety
More like this →
arXiv Machine Learning
1d ago

Certifying Compressed Language Models: An Audit and a Statistical Toolkit

arXiv:2608. 15046v1 Announce Type: new Abstract: A fraction of a point of benchmark accuracy is the usual evidence that a compressed model is equivalent to its original.

By Amogh Singh
llmsbenchmarks
More like this →