arXiv AI By Yang Liu, Bin Chong, Chongyang Zhang, Hao Zheng, Jiayu Liang, Xu Kefu

Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching

Read the original on arXiv AI →

arXiv:2608. 12331v1 Announce Type: cross Abstract: Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory bottleneck during decoding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 7

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV is a training‑free key‑value cache compression technique for Large Reasoning Models that uses beacon queries—compact representatives of query clusters—to predict which KV pairs will be revisited during long‑horizon reasoning. By focusing on Thought Revisiting Tokens that re‑attend distant context, BeaconKV reduces memory usage up to 5.8× and improves throughput by over 4.3× while largely preserving cache accuracy across multiple open‑source LRMs and reasoning benchmarks.

By Janghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook Choi
arXiv Machine Learning
4d ago

EpiKV: Epiphany-Aware KV Cache Eviction Without the Attention Matrix

The paper introduces EpiKV, an epiphany‑aware key–value cache eviction strategy that avoids using the attention matrix. It leverages hidden‑state shifts and recent query–key relevance to rank cached tokens, matching or surpassing the performance of existing attention‑based eviction methods while remaining compatible with fast inference kernels. Experiments on multiple benchmarks show that EpiKV improves inference throughput without sacrificing accuracy.

By Steven Kolawole, Virginia Smith