SPHQuant introduces a rotation‑free spherical weight‑only quantization framework for Vision‑Language Models, decomposing 8‑dimensional weight vectors into sign, radius, and a positive unit direction. By isolating outlier magnitudes in the radius and allocating extra precision there, it mitigates accuracy loss at extreme low bit‑widths. The method also employs a compact positive‑direction codebook with angular fine‑tuning and a hardware‑friendly GEMV kernel, achieving state‑of‑the‑art performance while boosting decode throughput by 30.3% on RTX A6000 compared to QTIP.
By Kewei Zhang, Zheng Chen, Haotong Qin, Yulun Zhang
Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones mak...
arXiv:2608. 04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step.
By Samuel Fern\'andez-Mendui\~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr
arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.
By Donghyun Son, Euntae Choi, Sungjoo Yoo
arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.
By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung
arXiv:2606. 15652v1 Announce Type: new Abstract: 4-bit quantization significantly reduces the memory footprint and accelerates the inference of large language models (LLMs).
By Yangjia Hu, Haodong Wang, Zicong Hong, Qianli Liu, Quanxin Shou, Jian Lin, Song Guo, Xiaowei Shen, Xiangjun Huang, Dian Wang, Jian Yang
arXiv:2606. 07116v1 Announce Type: cross Abstract: Low-bit quantization has been widely adopted to accelerate the inference of large language models (LLMs) by significantly reducing computational cost and memory usage.
By Haoqi Wang, Lorenz K. Mueller, Jiawei Zhuang, Mathieu Salzmann, Lukas Cavigelli
ConQuR introduces a lightweight post‑training rotation calibration for large language model activation quantization. By learning orthogonal rotations that align normalized activations with the corners of an inscribed hypercube, the method distributes activation energy evenly and can be updated online without storing activations. Experiments on Llama‑2 and Llama‑3 models (3B–70B) show competitive or improved perplexity and reasoning performance while avoiding costly training or large offline storage.
By Chayne Thrash, Ali Abbasi, Soheil Kolouri
arXiv:2606. 03458v1 Announce Type: new Abstract: Test-time scaling is a powerful approach to obtain better reasoning in large language models, but it becomes memory-bottlenecked during long-horizon decoding, as the KV-cache grows.
By Lorenz K. Muller, Philippe Bich, Chiara Boretti, Hyun-Min Chang, Jiawei Zhuang, Lukas Cavigelli
arXiv:2608.30384v1 Announce Type: new
Abstract: By introducing RSLM (Rotated Scaled Lloyd-Max), a family of training-free vector quantization codecs compressing embeddings to 1--4 bits per dimension,...
By Rastislav Lenhardt, Teodora Dobos, Thomas Vecchiato, Jiri Isa, Igor Ginzburg
The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.
By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu
arXiv:2608. 02691v1 Announce Type: cross Abstract: The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-low-bit quantization increasingly important.
By Vincent-Daniel Yun, Woosang Lim, Minsoo Cheong, Sunwoo Lee, Murali Annavaram, Sai Praneeth Karimireddy, Sungjoo Yoo