arXiv Machine Learning By Lin Zhang

SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV Retrieval

Read the original on arXiv Machine Learning →

arXiv:2608. 03228v1 Announce Type: new Abstract: Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.