← Back to all news
arXiv AI October 7, 2026 By Sara Abdali, Jongwoo Ko, Pashmina Cameron

AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • agents
  • efficiency
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv Machine Learning
Aug 26

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

arXiv:2608.23843v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compressio...

By Zizhong Wang, Jieying Wang, Zhao Zhang, Jiajia Li
llmsefficiencybenchmarks
More like this →
arXiv Machine Learning
3d ago

SlimKV: Joint Token-Feature KV Cache Compression with Reconstruction-Free Beacon Attention

arXiv:2610.02953v1 Announce Type: new Abstract: Long-context LLM serving is increasingly bottlenecked by KV-cache memory, especially in resource-constrained scenarios. Among existing KV-cache compres...

By Zihan Teng, Jiayu Zhao, Wentao Ren, Minhao Fan, Tianrui Ma, Song Chen, Weichen Liu
llmsragefficiency
More like this →
arXiv AI
Jul 21

SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation

arXiv:2607. 16213v1 Announce Type: new Abstract: Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with context length, creating a major bottleneck.

By Soumia Bouyahiaoui, Manel Kara laouar, Aicha Boutorh, Mohamed Hadj Ameur
llmsefficiencysafety
More like this →
arXiv AI
Sep 10

Attention-Weighted Value Projection for KV-Cache Compression

arXiv:2604.11501v2 Announce Type: replace-cross Abstract: Rank reduction discards dimensions; quantization keeps them at lower precision. Comparing the two requires a choice of what compression shoul...

By Samuel Salfati
efficiency
More like this →
arXiv Machine Learning
Jul 28

DynaCalKV: Key-Value Cache Compression via Head Grouping and Adaptive Rank Allocation

arXiv:2607. 24331v1 Announce Type: new Abstract: As the inference phase of Large Language Models (LLMs) requires handling long context windows, the Key-Value (KV) cache initially appears to address this challenge but eventually becomes a significant bottleneck as the context window continues to grow.

By Tan T. Nguyen, Quan V. Dang
llmsefficiencysafety
More like this →
arXiv Machine Learning
Jun 29

Learning to Evict from Key-Value Cache

arXiv:2602. 10238v2 Announce Type: replace-cross Abstract: The growing size of Large Language Models (LLMs) makes efficient inference challenging, primarily due to the memory demands of the autoregressive Key-Value (KV) cache.

By Luca Moschella, Laura Manduchi, Ozan Sener
llmsagentsreinforcement-learningefficiencybenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea