arXiv AI By Mukund Agarwalla, Chih-Jen Lin

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

Read the original on arXiv AI →

TopK-Guided is a training‑free method that improves activation sparsity for large language model inference by combining token‑level sparsity adaptation with block‑level budget allocation that accounts for block sensitivity. It addresses limitations of existing methods like TEAL, which adapts sparsity per token but lacks tight control, and WINA, which enforces a fixed sparsity across all tokens and blocks. Experiments on Llama‑2 and Llama‑3 show that TopK‑Guided consistently yields better perplexity and downstream accuracy while maintaining similar compute costs to WINA, especially at high sparsity levels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 23

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV introduces a compensation‑aware sparse attention framework for long‑context LLM inference. It partitions tokens into blocks and optimizes token selection to minimize the error introduced by block‑level mean compensation, using compact block‑level statistics. Experiments on RULER and LongBench‑Pro show CompKV outperforms other sparse baselines and achieves up to a 6.85× speedup over full attention.

By Zhen Huang, Ruizhe Yao, Danyi Liu, Xinrui Chen, Shuwei Li, Siru Zhong, Zijian Cao, Yushan Lai, Mingming Guo, Weijie Zheng, Haohuan Fu
arXiv Machine Learning
Aug 10

The Sparsity Whisperer

arXiv:2608. 06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs.

By Linghao Kong, Inimai Subramanian, Micah Adler, Dan Alistarh, Dan Gutfreund, Nir Shavit