arXiv:2608. 06291v1 Announce Type: cross Abstract: We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of the Hessian.
By Johann Birnick, Rayan Saab
arXiv:2607. 07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining.
By Donghyun Lee, Yuhang Li, Ruokai Yin, Priyadarshini Panda
arXiv:2605. 13768v2 Announce Type: replace-cross Abstract: This is the second part of the work investigating quantized matrix multiplication (MatMul).
By Or Ordentlich, Yury Polyanskiy
arXiv:2606. 00542v1 Announce Type: new Abstract: Shampoo-style optimizers approximate gradient covariance matrices using Kronecker-factored structures.
By Bing Liu, Wenjie Zhou, Chengcheng Zhao
arXiv:2505. 22988v3 Announce Type: replace-cross Abstract: The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible.
By Albert Tseng, Zhaofeng Sun, Christopher De Sa
The paper introduces a scalable Kronecker-based approximation that captures cross-layer interactions without storing the full Fisher matrix, making Hessian analysis feasible for billion-parameter language models. It identifies consistent vulnerability patterns, notably that value projection layers are the most sensitive and exhibit strong cross-layer correlations across various model families. Experiments on quantization, sparsification, inter-layer corruption, and fine-tuning show that the approximation correlates strongly with performance degradation and recovery, providing a practical tool for identifying fragile components and guiding compression and optimization strategies.
By Viacheslav Yusupov, Daria Cherniuk, Evgeny Frolov
arXiv:2601. 21626v2 Announce Type: replace-cross Abstract: Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing quantization error.
By Jinhao Zhang, Yunquan Zhang, Zicheng yan, Boyang Zhang, Jun Sun, Daning Cheng
arXiv:2609.37416v1 Announce Type: new
Abstract: Post-training quantization (PTQ) methods in the GPTQ family minimize a layer-wise reconstruction error on a uniform grid whose scale must be chosen; th...
By Jonas von Berg, Massimiliano Datres, Carlo Knei{\ss}l, Gitta Kutyniok
arXiv:2605.11222v2 Announce Type: replace
Abstract: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (...
By Ryan Lucas, Mehdi Makni, Xiang Meng, Adam Deng, Rahul Mazumder
arXiv:2606. 30523v1 Announce Type: new Abstract: Covariance matrices serve as compact descriptors of feature distributions in many machine-learning pipelines, including domain adaptation and Gaussian embeddings.
By Woojoo Na, Jennifer Dy
arXiv:2603. 04956v2 Announce Type: replace Abstract: This paper considers the problem of converting a given dense linear layer to low precision.
By Egor Lifar, Semyon Savkin, Or Ordentlich, Yury Polyanskiy
arXiv:2606. 01412v1 Announce Type: new Abstract: Post-training quantization is widely used for compressing large neural networks, but aggressive low-bit quantization can significantly degrade model quality.
By Shihao Zhang, Rayan Saab