arXiv AI By Bole Ma, Jan Eitzinger, Harald K\"ostler, Gerhard Wellein

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics

Read the original on arXiv AI →

arXiv:2606. 01502v1 Announce Type: cross Abstract: Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit is now a small, reusable chunk.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

VestigeKV: The NoPE-MLA KV Cache Carries Its Own Eviction Signal in a Vestigial Branch

VestigeKV is a new KV‑cache technique that uses a 64‑dimensional vestigial branch—originally a RoPE component repurposed during NoPE training—as a query‑independent eviction signal. By reading only 11 % of each cache row, the method partitions the cache into an attended tier (top‑m rows) and an archive tier (all other rows), which is GPU‑resident and never deleted. The approach achieves near‑perfect retrieval (1.00 at 8×, 0.92 at 32×) without any training, quantization, or changes to weights or kernels.

By WenJie Fan