Hugging Face Blog

Overview of natively supported quantization schemes in 🤗 Transformers

arXiv Machine Learning
Jul 24

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

arXiv:2607. 21446v1 Announce Type: new Abstract: Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent.

By Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann
arXiv AI
Sep 10

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

The paper presents a second‑order theory of quantization noise for matrix multiplication, characterizing quantization formats by the variance they assign to each element. It derives a closed‑form signal‑to‑noise‑ratio law for floating‑point rounding, introduces an upper bound κ* that cannot be exceeded by any function‑preserving linear transform, and proposes KBBQ—a method that parameterizes how closely a transform approaches this bound. Experiments on W4A4 across four base models and two FP4 formats show that KBBQ outperforms the previous state of the art without extra deployment‑time computation.

By Lexington Whalen, Yuki Ito, Ryo Sakamoto