arXiv AI By Johann Birnick, Rayan Saab

BaKron: Efficient Quantization with Kronecker-Factored Hessians

Read the original on arXiv AI →

arXiv:2608. 06291v1 Announce Type: cross Abstract: We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 4

Model-Preserving Adaptive Rounding

arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.

By Albert Tseng, Zhaofeng Sun, Christopher De Sa
arXiv Machine Learning
Jul 30

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding

arXiv:2607. 27042v1 Announce Type: cross Abstract: Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric.

By Jiale Chen, Torsten Hoefler, Dan Alistarh