arXiv Machine Learning

PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration

arXiv:2502. 00527v2 Announce Type: replace Abstract: The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability.

arXiv Computer Vision
Sep 22

SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.

By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
arXiv Machine Learning
1d ago

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs

ConQuR introduces a lightweight post‑training rotation calibration for large language model activation quantization. By learning orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, the method distributes activation energy evenly and can be updated online without storing activations. Experiments on Llama‑2 and Llama‑3 models (3B–70B) show competitive or improved perplexity and reasoning performance while avoiding costly training or large offline storage.

By Chayne Thrash, Ali Abbasi, Soheil Kolouri
arXiv Computation and Language
Sep 18

D-Quant: Driftable Entropy Coding for KV Cache Quantization

The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.

By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu