arXiv AI By Johann Birnick, Rayan Saab

BaKron: Efficient Quantization with Kronecker-Factored Hessians

Read the original on arXiv AI →

arXiv:2608. 06291v1 Announce Type: cross Abstract: We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv AI
Sep 15

WaterKron and FlipFlop Hessian: Information-Theoretically Grounded Quantization with Kronecker-factored Hessians

The paper introduces WaterKron, a method that integrates two-sided GPTQ with row- and column-dependent waterfilling scales and entropy coding for post‑training quantization. It derives a high‑rate distortion measure relative to the full Hessian, introducing a Kronecker‑Hessian mismatch factor Φ that quantifies the distortion penalty of using a Kronecker approximation. Minimizing Φ leads to a Gaussian covariance‑fitting problem solved via classical flip‑flop updates, yielding a FlipFlop Hessian that empirically improves KL divergence and perplexity compared to other Hessian choices.

By Johann Birnick, Rayan Saab
arXiv Machine Learning
Jul 30

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

arXiv:2607. 27042v1 Announce Type: cross Abstract: Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric.

By Jiale Chen, Torsten Hoefler, Dan Alistarh