← Back to all news
arXiv Machine Learning June 15, 2026 By Wei Jiang, Wei Wang

Sub-Token Routing for KV Cache Compression

Read the original on arXiv Machine Learning →

arXiv:2604. 21335v3 Announce Type: replace Abstract: Transformer inference often requires a large KV cache, especially for long-context language modeling and multimodal generation.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

  • llms
  • efficiency
  • multimodal

Related stories

arXiv AI
Jun 16

PolyKV: Heterogeneous Retention and Allocation for KV Cache Compression

arXiv:2606. 15157v1 Announce Type: cross Abstract: KV cache compression is essential for reducing the memory cost of long-context large language model inference.

By Chao Fei, Panos Kalnis
llmsefficiency
More like this →
arXiv Machine Learning
Jul 31

Back from the Future: Key-Value Cache Management by Counter-Causal Surprise

arXiv:2607. 27600v1 Announce Type: new Abstract: Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years.

By Stephen Gould, Anton van den Hengel
llmsefficiencymultimodalbenchmarks
More like this →
arXiv AI
Jul 3

Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

arXiv:2607. 01237v1 Announce Type: cross Abstract: Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited throughput.

By Shen Han, Yuyang Wu
llmsefficiency
More like this →
arXiv AI
Jun 17

AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor

arXiv:2606. 17872v1 Announce Type: cross Abstract: Large language models (LLMs) outperform earlier architectures on generative inference and long-context tasks, but their large size introduces significant challenges in memory usage, energy cost, and on-device deployment.

By Ning Ni, Yingjie Lao
llmsefficiencysafety
More like this →
arXiv AI
Jul 7

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression

arXiv:2607. 01237v2 Announce Type: replace-cross Abstract: Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and incurs high decoding latency and limited throughput.

By Shen Han, Yuyang Wu, Junpu Yu, Olexandr Isayev
llmsefficiency
More like this →
arXiv Machine Learning
Aug 5

AnchorKV: Anchor-Residual KV Cache Compression

arXiv:2608. 02901v1 Announce Type: new Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference.

By Malik Khalaf, Yara Shamshoum, Nitzan Hodos, Yuval Sieradzki, Assaf Schuster
llmsefficiency
More like this →