← Back to all news
Hugging Face Blog May 16, 2024

Unlocking Longer Generation with Key-Value Cache Quantization

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jul 9

Fractal KV-Cache Archives: Lossless Symbolic Storage with In-Place Retrieval for Long-Context LLM Inference

arXiv:2607. 07144v1 Announce Type: new Abstract: The key-value (KV) cache dominates the memory cost of long-context autoregressive inference, and a growing body of work compresses it through quantization, eviction, or offloading.

By Vladimir Gusev
llmsefficiency
More like this →
arXiv Machine Learning
Aug 5

AnchorKV: Anchor-Residual KV Cache Compression

arXiv:2608. 02901v1 Announce Type: new Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference.

By Malik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster
llmsefficiency
More like this →
arXiv AI
Jul 16

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.

By Donghyun Son, Euntae Choi, Sungjoo Yoo
llmsefficiencysafety
More like this →
arXiv AI
Jun 6

Channel-Wise Mixed-Precision Quantization for Large Language Models

arXiv:2410. 13056v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated remarkable success across a wide range of language tasks, but their deployment on edge devices remains challenging due to the substantial memory requirements imposed by their large parameter sizes.

By Zihan Chen, Bike Xie, Jundong Li, Cong Shen
llmsefficiency
More like this →
arXiv Machine Learning
Aug 17

KV Cache Compression Through the Lens of Transform Coding

arXiv:2608. 14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference.

By Hannah Laus, Claudio Mayrink Verdun, Hao Wang, Flavio du Pin Calmon, Felix Krahmer
llmsefficiencybenchmarks
More like this →
arXiv Machine Learning
Jun 29

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

arXiv:2605. 06675v2 Announce Type: replace Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottleneck for serving.

By Fei Zuo, Zikang Zhou, Hao Cong, Xiaoyan Xi, Ho Fai Leung
llmsefficiency
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea