arXiv Machine Learning By Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Read the original on arXiv Machine Learning →

CommunityKV is a new framework that treats sparse attention as a community detection problem, building a token graph from $QK^T$ scores and partitioning it into semantically coherent communities. It updates token communities in constant time during streaming decoding, avoiding costly global re‑partitioning. Experiments on Qwen3 and Llama‑3.1 show that CommunityKV can increase end‑to‑end generation throughput by up to 1.25×, and with query‑group graph aggregation up to 1.71×, while maintaining comparable accuracy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 24

CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference

arXiv:2606. 24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware.

By Xiaolin Lin, Jingcun Wang, Olga Kondrateva, Yiyu Shi, Bing Li, Grace Li Zhang