arXiv:2601. 03043v4 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, placing significant demands on inference efficiency.
By Junhao Hu, Fangze Li, Mingtao Xu, Feifan Meng, Shiju Zhao, Tiancheng Hu, Ting Peng, Anmin Liu, Wenrui Huang, Chenxu Liu, Ziyue Hua, Tao Xie
The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.
By Zizhuo Fu, Runsheng Wang, Meng Li
The paper introduces SCSP, a training‑free framework that improves long‑context embeddings by selectively pooling informative tokens. SCSP partitions documents into sentence‑aware chunks, adds a semantic compression prompt to each chunk, and uses prompt‑isolated attention masks to estimate token importance. The selected tokens’ intermediate‑layer representations are aggregated to form the final embedding, yielding consistent performance gains across zero‑shot and fine‑tuned models on long‑context benchmarks.
By Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu
arXiv:2603. 05353v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) for long-context question answering is bottlenecked by inference-time prefilling over large retrieved contexts.
By Xin Teng, Canyu Zhang, Shaoyi Zheng, Danyang Zhuo, Tianyi Zhou, Shenji Wan
arXiv:2606. 10944v1 Announce Type: new Abstract: We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees.
By Albert Gong, Annabelle Michael Carrell, Raaz Dwivedi, Lester Mackey
arXiv:2608. 13578v1 Announce Type: cross Abstract: Transformer architectures rely on dense self-attention to model long-range dependencies, but this mechanism exhibits quadratic complexity with respect to sequence length.
By Rachid Arezki