arXiv Machine Learning

Why Do Accumulated Transformations Extrapolate?

arXiv:2606. 24975v1 Announce Type: new Abstract: PaTH Attention showed that replacing RoPE's position-indexed rotations with accumulated data-dependent Householder reflections yields strong length extrapolation, though performance degrades at extreme context lengths.

arXiv Machine Learning
Jun 24

Selective Rotary Position Embedding

arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.

By Sajad Movahedi, Timur Carstensen, Arshia Afzal, Frank Hutter, Antonio Orvieto, Volkan Cevher
arXiv Machine Learning
Sep 10

Content-Based Addressing for Long Context

The paper proposes a content‑based addressing scheme for long‑context models that replaces the growing token counter in Rotary Position Embedding (RoPE) with unit‑level addresses derived from the content of each unit. By dividing the token stream into units, the method preserves local RoPE behavior while allowing new units to be addressed via learned content maps, avoiding positional mismatches when extending context length. Experiments on character‑level Tiny Shakespeare show that a model trained on 256‑character contexts achieves lower perplexity at 4096 characters using this scheme, and a second diagnostic demonstrates retrieval of multiple serialized facts.

By Mahesh Godavarti
arXiv Computation and Language
Sep 11

Distance generalization in transformers: why bother with positional encoding?

The paper investigates distance generalization in transformer models, focusing on how well they can handle changes in inter-token distances between training and inference while keeping context length fixed. Using two synthetic delay-copy tasks that require copying tokens after finite delays, the authors evaluate the impact of positional encoding schemes (RoPE, ALiBi, and NoPE), the diversity of distances seen during training, and the conditions under which distance transfer learning is beneficial or detrimental. Their comprehensive study highlights the importance of understanding the underlying mechanisms that govern distance generalization in transformers.

By Daniel Henrik Nevermann, Claudius Gros