arXiv AI By Samuel Fern\'andez-Mendui\~na, Amir Ziashahabi, Eduardo Pavez, Antonio Ortega, Salman Avestimehr

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

Read the original on arXiv AI →

arXiv:2608. 04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.