arXiv AI
Jul 3

Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning

arXiv:2607. 01241v1 Announce Type: cross Abstract: Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is often spread across multiple locations and connected through both local syntactic dependencies and global semantic relations.

By Yaxin Gao, Yao Lu, Jinhong Deng, Jiaqi Nie, Zhe Tang, Jian Zhang, Zhaowei Zhu, Shanqing Yu, Qi Xuan, Joey Tianyi Zhou
arXiv Machine Learning
1d ago

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

CommunityKV is a new framework that treats sparse attention as a community detection problem, building a token graph from $QK^T$ scores and partitioning it into semantically coherent communities. It updates token communities in constant time during streaming decoding, avoiding costly global re‑partitioning. Experiments on Qwen3 and Llama‑3.1 show that CommunityKV can increase end‑to‑end generation throughput by up to 1.25×, and with query‑group graph aggregation up to 1.71×, while maintaining comparable accuracy.

By Joe McKenna, Anastasios Alexandridis, Nathan Susanj, Jing Liu