arXiv AI

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

GrowPage is an on‑demand key–value (KV) budgeting framework designed to improve large language model (LLM) reasoning serving. It treats KV capacity as a runtime resource, using lightweight dual‑timescale query summaries to track recent and long‑term attention patterns and estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page, integrating with PagedAttention’s page‑level memory abstraction to maintain continuous batching and prefix caching.

arXiv Machine Learning
Sep 7

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV is a training‑free key‑value cache compression technique for Large Reasoning Models that uses beacon queries—compact representatives of query clusters—to predict which KV pairs will be revisited during long‑horizon reasoning. By focusing on Thought Revisiting Tokens that re‑attend distant context, BeaconKV reduces memory usage up to 5.8× and improves throughput by over 4.3× while largely preserving cache accuracy across multiple open‑source LRMs and reasoning benchmarks.

By Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
arXiv AI
3d ago

ActKV: Efficient LLM Agents through Action-Guided KV Cache Management

ActKV is a new KV cache compression framework designed for agentic large language model (LLM) inference. It prioritizes cache entries that contribute to action generation, using action-oriented eviction, confidence-driven budget allocation, and page-aware compression to reduce memory usage while preserving accuracy. In long-trace tasks, ActKV retains 98.53% of FullKV’s accuracy using only 25.98% of its peak memory and boosts token and task throughput by 3.97× and 3.58×, respectively.

By Zihan Wang, Cheng Tang, Lei Gong, Chao Wang, Wenqi Lou, Teng Wang, Xuehai Zhou
arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang