KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 08382v1 Announce Type: cross Abstract: Low-rank projection has emerged as a promising approach for compressing the KV cache by exploiting hidden-dimension redundancy.
arXiv:2605. 06675v2 Announce Type: replace Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottleneck for serving.
arXiv:2604.11501v2 Announce Type: replace-cross Abstract: Rank reduction discards dimensions; quantization keeps them at lower precision. Comparing the two requires a choice of what compression shoul...
The paper introduces D-Quant, a KV cache quantization framework that addresses the memory bottleneck of large language models by using a drift mechanism to convert entropy-coded representations into fixed-size bitstreams. This approach leverages the non-uniform distribution of KV cache values—after rotation and normalization, they approximate a normal distribution—allowing entropy coding to assign shorter codewords to frequent symbols while maintaining regular memory layouts suitable for parallel attention kernels. D-Quant thus aims to reduce memory footprint and bandwidth usage without sacrificing performance.
arXiv:2608. 04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step.
arXiv:2606. 24033v1 Announce Type: new Abstract: Existing low-bit KV-cache quantizers often treat each cached key as a flat vector.