arXiv AI By Yuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao Diao

OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference

Read the original on arXiv AI →

arXiv:2510. 07651v3 Announce Type: replace-cross Abstract: Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all key-value (KV) states scales linearly with sequence length and batch size.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.