arXiv Machine Learning By Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

Read the original on arXiv Machine Learning →

arXiv:2608. 01662v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Faster Than Flash: Exploiting Attention Sparsity for Efficient Long-Context Decoding

The paper introduces Faster Flash Decoding (FFD), a hardware‑algorithm co‑design that fuses the selector and computation into a single kernel to eliminate memory‑bandwidth bottlenecks in long‑context decoding. By replacing metadata indices with low‑bit quantized, content‑aware scanning and employing a top‑delta strategy for dynamic block filtering, FFD achieves up to 11.6× kernel‑level speedup and scales to 256K context length while preserving model accuracy. The approach is training‑free, plug‑and‑play, and demonstrates significant throughput gains on benchmarks such as RULER and LongBench.

By Zhigeng Liu, Zhiyuan Ning, Ruixiao Li, Xiaoran Liu, Yuerong Song, Min Zhang, Ziwei He, Xipeng Qiu
arXiv Machine Learning
Sep 11

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention

HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention proposes a two-stage hierarchical indexer that replaces the flat token scan used in token-level sparse attention mechanisms like DeepSeek Sparse Attention. The method first performs block-level coarse filtering to discard irrelevant regions, then applies the original token-level indexer only within the retained candidate blocks, preserving the same top-sparse pattern for downstream attention. Benchmarks show HISA achieves significant speedups at 64K context and matches the quality of DeepSeek-V3.2 and GLM-5 without additional training.

By Yufei Xu, Fanxu Meng, Fan Jiang, Yuxuan Wang, Ruijie Zhou, Zhaohui Wang, Jiexi Wu, Zhixin Pan, Xiaojuan Tang, Wenjie Pei, Tongxuan Liu, Di Yin, Xing Sun, Muhan Zhang
arXiv Computation and Language
Sep 23

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

arXiv:2609.26368v1 Announce Type: new Abstract: Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context dem...

By Jianyu Wei, Yizhao Gao, Qihao Zhang, Shimao Chen, Zhengju Tang, Yu Cheng, Shengjie Zhou, Zihan Jiang, Yifan Song, Hailin Zhang, Liang Zhao, Bo Yang, Gang Wang, Shijie Cao, Fuli Luo