arXiv AI By Zhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu

KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving

Read the original on arXiv AI →

arXiv:2605. 09735v2 Announce Type: replace-cross Abstract: Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS events arrive asynchronously, and logical histories fragment over time.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.