arXiv Machine Learning By Lin Zhang

SAKI: Score-Aware Low-Rank Key Indexing with Random-Matrix Noise Correction for KV Retrieval

Read the original on arXiv Machine Learning →

arXiv:2608. 03228v2 Announce Type: replace Abstract: Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during inference.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.