arXiv AI

HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning

arXiv:2601. 21626v2 Announce Type: replace-cross Abstract: Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomenon because it focuses solely on minimizing quantization error.

arXiv Machine Learning
Sep 17

Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate

The paper introduces DASH-Q, a post‑training quantization method that uses a diagonal Hessian approximation and iterative weighted least squares to reduce noise in curvature estimates. By discarding noisy cross‑channel dependencies, DASH‑Q preserves salient feature power and achieves superior performance in ultra low‑bit quantization. Across five large language models, it improves zero‑shot accuracy by an average of 7.01% and up to 14.01% over the strongest baselines, even with very small calibration datasets.

By Jaemin Kim, Sungkyun Kim, Junyeol Lee, Jiwon Seo
Hugging Face Trending Papers
Aug 12

HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression

Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance but are hindered by high computational complexity and encoding-decoding mismatches across heterogeneous hardware platforms. Uniform fixed-precision quantization alleviates these issues but suffers severe quality degradation at low bit widths because it ignores differences in the quantization sensitivities of individual layers.

arXiv AI
Aug 26

Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM Compression

The paper introduces the "Compression Trinity," a unified framework that jointly applies sparsity, quantization, and low‑rank approximations to compress large language models. It presents several methods—MKOR, SLoPe, OPTIMA, PATCH, and SLiM—that leverage these three pillars to accelerate training, reduce memory bandwidth, and recover accuracy, achieving significant speedups and accuracy gains over existing techniques. The results demonstrate that combining all three compression strategies is essential for efficient, scalable, high‑performance LLM deployment.

By Mohammad Mozaffari