Overview of natively supported quantization schemes in 🤗 Transformers
Related stories
A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.
WUSH: Near-Optimal Adaptive Transforms for LLM Quantization
arXiv:2512. 00956v3 Announce Type: replace Abstract: Quantizing LLM weights and activations is a standard approach for efficient deployment, but a few extreme outliers can stretch the dynamic range and amplify low-bit quantization errors.
Introducing AutoRound: Intel’s Advanced Quantization for LLMs and VLMs
Exploring Quantization Backends in Diffusers
Quanto: a PyTorch quantization backend for Optimum
On the Expressive Power of Transformers
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
VQ-bench: A Composable Vector Quantization Framework
arXiv:2608. 11240v1 Announce Type: new Abstract: Vector quantization is an old problem but has recently become central to AI infrastructure.
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck,...
KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization
The paper presents a second‑order theory of quantization noise for matrix multiplication, characterizing quantization formats by the variance they assign to each element. It derives a closed‑form signal‑to‑noise‑ratio law for floating‑point rounding, introduces an upper bound κ* that cannot be exceeded by any function‑preserving linear transform, and proposes KBBQ—a method that parameterizes how closely a transform approaches this bound. Experiments on W4A4 across four base models and two FP4 formats show that KBBQ outperforms the previous state of the art without extra deployment‑time computation.