SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.
By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones mak...
arXiv:2608.30384v1 Announce Type: new
Abstract: By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension,...
By Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato, Jiri Isa, Igor Ginzburg
arXiv:2502. 00527v2 Announce Type: replace Abstract: The KV cache in large language models is a dominant factor in memory usage, limiting their broader applicability.
By Songhao Wu, Ang Lv, Xiao Feng, Yufei Zhang, Xun Zhang, Guojun Yin, Wei Lin, Rui Yan
arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.
By Donghyun Son, Euntae Choi, Sungjoo Yoo
arXiv:2608. 04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step.
By Samuel Fern\'andez-Mendui\~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr
ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.
By Mehdi Makni, Ryan Lucas, Rahul Mazumder
The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.
By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu
arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.
By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung
PRQuant introduces a training‑free, low‑overhead method for low‑bit quantization of linear layers by permuting input channels that cause the largest quantization error into contiguous tail blocks and precomputing residual weight sub‑tensors. The approach eliminates the need for online gathering during inference, converting scattered residual compensation into a regular tail‑augmented GEMM and thereby reducing latency. Experiments show that PRQuant lowers down‑projection reconstruction error and outperforms standard MXFP4 and other post‑training quantization baselines on five downstream benchmarks, improving accuracy by up to 1.24 points on Qwen3‑4B‑Instruct‑2507.
By Peiran Wang, Anqi Wang, Jiaying Zhao, Huiwen Yang, Zhenyu Ming, Rongqian Wang, Yiwu Yao, Kun Tian, Xin Yao, Gong Zhang, Fan Yang, Zhongyi Huang
arXiv:2510. 00566v4 Announce Type: replace-cross Abstract: Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate verification, making it the primary bottleneck in the search process.
By Vansh Ramani, Alexis Schlomer, Akash Nayar, Sayan Ranu, Jignesh M. Patel, Panagiotis Karras
Vision Foundation Models (VFMs) have significantly advanced dense feature matching, yet severe in-plane rotation remains a critical challenge. Existing solutions face a fundamental dilemma: data-driven methods require inefficient parameter scaling to implicitly learn rotations, whereas strictly equivariant networks lack the semantic capacity of modern VFMs.