arXiv Machine Learning By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache

Read the original on arXiv Machine Learning →

arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.