arXiv AI By Gradwell Dzikanyanga, Yanqi Pan, Weihao Yang, Donglei Wu, Wen Xia, Hao Huang

High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration

Read the original on arXiv AI →

arXiv:2607. 16248v1 Announce Type: cross Abstract: Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and bandwidth overheads.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.