arXiv AI By Shaowen Wang, Yuke Zheng, Tansheng Zhu, Shuang Chen, Shaofan Liu, Suncong Zheng, Jian Li

AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally

Read the original on arXiv AI →

arXiv:2607. 19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design

The paper proposes a principled way to design hybrid transformer architectures that combine Full Attention (FA) and Linear Attention (LA). By introducing two intervention metrics—RoPE Frequency Importance Score (RFIS) and RoPE Positional Dependence (RPD)—the authors identify a clear taxonomy of retrieval and positional heads, defining a Global Positional Band (GPBand) that aligns with training-length positional scales. Using these insights, they build a Head‑wise Hybrid Architecture (HwH) that assigns FA to global retrieval and LA to local positional modeling, achieving strong language modeling, improved retrieval, and superior zero‑shot long‑context extrapolation compared to standard Transformers and other hybrids.

By Runlin Shi, Bojian Yin, Guoqi Li
arXiv Computation and Language
3d ago

Scaling Parameter and Context in Attention: Native Sparse Attention from Mixture-of-Head

The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.

By Zizhuo Fu, Runsheng Wang, Meng Li
arXiv Machine Learning
Sep 10

Content-Based Addressing for Long Context

The paper proposes a content‑based addressing scheme for long‑context models that replaces the growing token counter in Rotary Position Embedding (RoPE) with unit‑level addresses derived from the content of each unit. By dividing the token stream into units, the method preserves local RoPE behavior while allowing new units to be addressed via learned content maps, avoiding positional mismatches when extending context length. Experiments on character‑level Tiny Shakespeare show that a model trained on 256‑character contexts achieves lower perplexity at 4096 characters using this scheme, and a second diagnostic demonstrates retrieval of multiple serialized facts.

By Mahesh Godavarti