← Back to all news
arXiv Machine Learning September 30, 2026 By Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • efficiency

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 17

KV Cache Compression Through the Lens of Transform Coding

arXiv:2608. 14191v1 Announce Type: new Abstract: The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference.

By Hannah Laus, Claudio Mayrink Verdun, Hao Wang, Flavio du Pin Calmon, Felix Krahmer
llmsefficiencybenchmarks
More like this →
arXiv AI
Jul 16

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.

By Donghyun Son, Euntae Choi, Sungjoo Yoo
llmsefficiencysafety
More like this →
arXiv AI
Aug 6

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

arXiv:2608. 04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step.

By Samuel Fern\'andez-Mendui\~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr
llmsefficiency
More like this →
arXiv Machine Learning
Jul 2

GSRQ: Gain-Shape Residual Quantization for Sub-1-bit KV Cache

arXiv:2607. 01065v1 Announce Type: new Abstract: The deployment of Large Language Models (LLMs) with extended context windows is increasingly constrained by the linear growth of Key-Value (KV) cache memory.

By Soosung Kim, Minjae Park, Eui-Young Chung, Jaeyong Chung
llmsefficiencysafety
More like this →
arXiv Machine Learning
Jun 8

AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization

arXiv:2605. 08692v2 Announce Type: replace Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference.

By Beshr IslamBouli, David Jin
llmsefficiency
More like this →
arXiv Machine Learning
4d ago

QuantMLA: Function-Aligned Dual-Path Quantization for Low-Bit MLA KV Caching

arXiv:2609.36760v1 Announce Type: new Abstract: Multi-Head Latent Attention (MLA) enables expressive multi-head attention with compact caches for its content and decoupled RoPE paths, yet cache memor...

By Zunhai Su, Yuxuan Sun, Jianchao Tan, Tao Zhang, Ruihan Hu, Yuchen Xie, Xunliang Cai, Ngai Wong
efficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea