The paper proposes a principled way to design hybrid transformer architectures that combine Full Attention (FA) and Linear Attention (LA). By introducing two intervention metrics—RoPE Frequency Importance Score (RFIS) and RoPE Positional Dependence (RPD)—the authors identify a clear taxonomy of retrieval and positional heads, defining a Global Positional Band (GPBand) that aligns with training-length positional scales. Using these insights, they build a Head‑wise Hybrid Architecture (HwH) that assigns FA to global retrieval and LA to local positional modeling, achieving strong language modeling, improved retrieval, and superior zero‑shot long‑context extrapolation compared to standard Transformers and other hybrids.
By Runlin Shi, Bojian Yin, Guoqi Li
Positional encodings (PEs) are essential for Transformers. Yet designing effective PEs for non-Euclidean graphs remains challenging.
arXiv:2606. 25293v1 Announce Type: new Abstract: Positional encodings (PEs) are essential for Transformers.
By Yipeng Zhang, Zhongtian Sun, Pietro Li\`o, Kelin Xia
arXiv:2602. 18948v2 Announce Type: replace Abstract: Transformer models contain substantial internal redundancy arising from coordinate-dependent representations and continuous symmetries, in model space and in head space, respectively.
By J. Fran\c{c}ois, L. Ravera
arXiv:2609.37921v1 Announce Type: new
Abstract: What are the inductive biases of a Transformer architecture? Existing theory on how the forward pass shapes representations either considers whether Tr...
By Erkan Turan, Gaspard Abel, Maks Ovsjanikov
arXiv:2609.15975v1 Announce Type: cross
Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
By Shwai He, Haichao Zhang, Shen Yan