arXiv AI By Bole Ma, Jan Eitzinger, Harald K\"ostler, Gerhard Wellein

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics

Read the original on arXiv AI →

arXiv:2606. 01502v1 Announce Type: cross Abstract: Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: attention's unit is now a small, reusable chunk.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.