arXiv Machine Learning

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache

arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.