arXiv Machine Learning By Hyungmin Kim, Minsoo Kim, Hongseok Kim, Jungwook Choi

Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving

Read the original on arXiv Machine Learning →

arXiv:2606. 06302v2 Announce Type: replace Abstract: Multi-turn LLM serving accumulates dialogue history whose Key-Value (KV) cache grows with every turn and every user, quickly exceeding the model weights themselves and making memory -- not compute -- the binding constraint on throughput.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.