arXiv AI

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration

arXiv:2607. 16248v1 Announce Type: cross Abstract: Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads.