arXiv AI By Jinsong Shu, Chenyang Wu, Zhongle Xie, Baokun Wang, Lidan Shou

MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models

Read the original on arXiv AI →

arXiv:2607. 22586v1 Announce Type: new Abstract: Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows linearly with context length and becomes a major bottleneck due to the large number of visual tokens.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.