arXiv AI By Zimu Zhao

Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention

Read the original on arXiv AI →

arXiv:2608. 19203v1 Announce Type: cross Abstract: Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head

The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.

By Zizhuo Fu, Runsheng Wang, Meng Li
arXiv Machine Learning
Jun 16

Rethinking the Role of Efficient Attention in Hybrid Architectures

arXiv:2606. 15378v1 Announce Type: cross Abstract: Modern language models increasingly adopt hybrid architectures that combine full attention with efficient attention modules, such as sliding-window attention (SWA) and recurrent sequence mixers.

By Ziqing Qiao, Yinuo Xu, Chaojun Xiao, Zhou Su, Zihan Zhou, Yingfa Chen, Xiaoyue Xu, Xu Han, Zhiyuan Liu