The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.
By Yi Su, Hong Liu, Guanghua Yu, Jianchen Zhu
arXiv:2609.24298v1 Announce Type: new
Abstract: What limits KV-cache compression at extreme bit-rates? We argue that it is not the choice of compression scheme, but how its budget is allocated across...
By Sihyeon Ha, Jaeho Lee, Yo-Seb Jeon
arXiv:2608.28911v1 Announce Type: new
Abstract: The key-value (KV) cache is the dominant memory bottleneck of long-context large language model (LLM) inference, growing linearly with context length....
By Daeha Lee, Do-Hyung Kim, Jae-Hong Kim
arXiv:2607. 23047v1 Announce Type: cross Abstract: Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but existing methods solve the allocation for a single fixed memory budget.
By Ashitabh Misra, Madhav Agrawal, Arham Jain, Tarek Abdelzaher
arXiv:2606. 24033v1 Announce Type: new Abstract: Existing low-bit KV-cache quantizers often treat each cached key as a flat vector.
By Fengfeng Liang, Yuechen Zhang, Jiaya Jia
arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.
By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung