← Back to all news
arXiv Machine Learning September 30, 2026 By Nenad Banfic

HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms
  • efficiency
  • multimodal
  • benchmarks

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Jun 29

RateQuant: Optimal Mixed-Precision KV Cache Quantization via Rate-Distortion Theory

arXiv:2605. 06675v2 Announce Type: replace Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottleneck for serving.

By Fei Zuo, Zikang Zhou, Hao Cong, Xiaoyan Xi, Ho Fai Leung
llmsefficiency
More like this →
arXiv Computer Vision
Sep 22

GPLQ: A General, Practical, and Lightning QAT Method for Vision Transformers

arXiv:2506.11784v2 Announce Type: replace Abstract: Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-wid...

By Guang Liang, Xinyao Liu, Jianxin Wu
llmscomputer-visionefficiency
More like this →
arXiv Machine Learning
4d ago

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck,...

By Jiale Chen, Vage Egiazarian, Eldar Kurti\'c, Torsten Hoefler, Dan Alistarh
efficiency
More like this →
arXiv Computer Vision
Sep 23

QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for World Models and Video Generation

arXiv:2609.26425v1 Announce Type: new Abstract: KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficien...

By Jiaqi Zhao, Xiaobin Hu, Bo Yin, Junpeng Jiang, Miao Zhang, Shuicheng Yan
diffusionefficiencybenchmarks
More like this →
arXiv Machine Learning
Jul 28

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

arXiv:2607. 24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness.

By M M Asif Ferdous
llmsefficiencymultimodal
More like this →
arXiv AI
Jul 16

NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache

arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.

By Donghyun Son, Euntae Choi, Sungjoo Yoo
llmsefficiencysafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea