arXiv Machine Learning

Stochastic Rounding Increases Small Singular Values

arXiv:2606. 00312v1 Announce Type: cross Abstract: Over the past half-dozen years, stochastic rounding (SR) has regained significant attention as a quantization scheme for low-precision floating-point arithmetic, with applications spanning numerical analysis and modern machine learning systems.

arXiv AI
Sep 10

KBBQ: A Predictive Noise Law and the Limits of Spectrum Flattening in FP4 Quantization

The paper presents a second‑order theory of quantization noise for matrix multiplication, characterizing quantization formats by the variance they assign to each element. It derives a closed‑form signal‑to‑noise‑ratio law for floating‑point rounding, introduces an upper bound κ* that cannot be exceeded by any function‑preserving linear transform, and proposes KBBQ—a method that parameterizes how closely a transform approaches this bound. Experiments on W4A4 across four base models and two FP4 formats show that KBBQ outperforms the previous state of the art without extra deployment‑time computation.

By Lexington Whalen, Yuki Ito, Ryo Sakamoto
arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv Computation and Language
Sep 11

Structured Transforms for Low-Overhead Quantization of Language Models

The paper revisits Kashin‑decomposition‑based weight quantization for large language models, introducing an improved algorithm that uses a sign‑randomized Discrete Cosine Transform (DCT) instead of a dense random orthogonal matrix. This change reduces per‑iteration cost from ≠(N^2) to ≠(N log N) and, combined with a greedy alternating‑update scheme, guarantees the four‑peak distribution needed for stable 2‑bit clustering while eliminating the need for multi‑restart k‑means. The resulting JAX pipeline, when paired with OPTQ‑style error compensation and QuIP‑style incoherence preprocessing, competes with state‑of‑the‑art quantization methods on OPT, Llama‑2, and Pythia at 4‑bit per channel, and remains numerically stable under stress configurations that cause other methods to diverge.

By Daria Cherniuk, Alexander Rudikov, Boris Kashin, Ivan Oseledets