HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2605. 06675v2 Announce Type: replace Abstract: Large language models cache all previously computed key-value (KV) pairs during generation, and this KV cache grows linearly with sequence length, making it a primary memory bottleneck for serving.
arXiv:2506.11784v2 Announce Type: replace Abstract: Vision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-wid...
arXiv:2609.38121v1 Announce Type: new Abstract: KV cache memory and bandwidth costs grow with context length and batch size, which limits efficient long-context inference. To address this bottleneck,...
arXiv:2609.26425v1 Announce Type: new Abstract: KV cache memory has become a major deployment bottleneck for video generation and world models, which motivates low-bit quantization study for efficien...
arXiv:2607. 24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness.
arXiv:2505. 18231v3 Announce Type: replace-cross Abstract: Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to the large size of key-value (KV) cache.