arXiv AI By Donghyun Son, Euntae Choi, Sungjoo Yoo

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

Read the original on arXiv AI →

arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.