arXiv Machine Learning By Lukas Haverbeck, Carmen Amo Alonso, Andres Felipe Posada-Moreno, Sebastian Trimpe, Marco Pavone

The risk of KV cache compression

Read the original on arXiv Machine Learning →

arXiv:2607. 01520v1 Announce Type: new Abstract: Transformer inference on long sequences is expensive because softmax attention repeatedly reads from a large KV cache.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 8

FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference

arXiv:2607. 06519v1 Announce Type: new Abstract: Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the layer-specific evidence needed for retrieval and multi-step reasoning.

By Anna C\'ordoba, Adam Puente Tercero, Nerea Angulo Hijo, Mar Linares Tercero, Julia Barrientos, Ainhoa Miranda, Jes\'us Olivera
arXiv AI
4d ago

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

The paper investigates how chunked KV‑cache compression, which reduces memory and attention costs by grouping consecutive tokens, introduces a new positional coordinate called a token’s phase. It finds that models using this compression exhibit phase sensitivity: retrieval performance can vary by up to 40 percentage points depending on a token’s phase, creating periodic weak spots that average benchmark scores hide. By pretraining transformers with different compression designs and applying causal interventions, the authors show that attention components specialize asymmetrically for different phases, and theoretical analysis suggests gradient dynamics may favor such specialization.

By Xingyu Zhu (Luke), Pu (Luke), Yi, Ziheng Cheng, Ang Lv, Jing Liu, Lexing Ying, Yiyuan Ma, Xin Dong