arXiv Machine Learning By Yang Pengju

SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM Serving

Read the original on arXiv Machine Learning →

arXiv:2606. 08635v1 Announce Type: new Abstract: Prefill-decode (PD) disaggregation decouples prompt processing from token generation, but it also turns the key-value (KV) cache into a network payload.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.