How Local Mixing Encodes Relative Position in Global NoPE Attention
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
arXiv:2512. 14391v3 Announce Type: replace-cross Abstract: In-context learning is fundamental to modern Large Language Models (LLMs); however, prevailing architectures impose a rigid and fixed contextual structure by assigning linear or constant positional indices.
The paper proposes an encoder Transformer that explicitly separates semantic, absolute positional (AP), and relative positional (RP) information, restricting the masked‑language‑modeling objective to the semantic stream. This disentanglement reveals that the AP subspace collapses into a low‑frequency two‑dimensional manifold reflecting document structure, that attention heads specialize into structure‑ and semantic‑oriented groups with RP supporting only the latter, and that standard positional encodings fail to robustly encode macroscopic structure. The approach preserves positional encoding and improves performance on 49 out of 65 linguistic phenomena in the Flash‑Holmes probing benchmark.
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
The paper introduces NAMOH, a native sparse attention mechanism that activates only a subset of heads per token, allowing each head to attend to a limited subsequence of tokens. By scaling the number of heads while keeping the active heads per token fixed, the method shortens head histories and reduces key‑value access without increasing overall storage. Experiments demonstrate that NAMOH can outperform fully activated models with the same parameter count and enable more efficient long‑context inference than smaller dense models.
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.