arXiv AI By Ziyao Tang, Pengkun Jiao, Xinhang Chen, Wei Liu, Shiyong Li, Jingjing Chen

Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache Eviction

Read the original on arXiv AI →

arXiv:2602. 08585v2 Announce Type: replace-cross Abstract: Given the quadratic complexity of attention, KV cache eviction is vital to accelerate model inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving

GrowPage is an on‑demand key–value (KV) budgeting framework designed to improve large language model (LLM) reasoning serving. It treats KV capacity as a runtime resource, using lightweight dual‑timescale query summaries to track recent and long‑term attention patterns and estimate demand evolution. At each capacity boundary, GrowPage either compresses KV states within the current allocation or acquires an additional physical page, integrating with PagedAttention’s page‑level memory abstraction to maintain continuous batching and prefix caching.

By Qiankun Ma, Yanjiang Zhou, Zinan Xiong, Haofei Wang, Zhen Song, Yang Xiang, Ziyao Zhang, Hairong Zheng