arXiv Machine Learning By Xinyi Wu, Siyuan Liu, Ali Jadbabaie

How Data Shapes RoPE Frequency Usage: From Positional Scale Matching to Length Generalization

Read the original on arXiv Machine Learning →

arXiv:2607. 07678v1 Announce Type: new Abstract: Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

Content-Based Addressing for Long Context

The paper proposes a content‑based addressing scheme for long‑context models that replaces the growing token counter in Rotary Position Embedding (RoPE) with unit‑level addresses derived from the content of each unit. By dividing the token stream into units, the method preserves local RoPE behavior while allowing new units to be addressed via learned content maps, avoiding positional mismatches when extending context length. Experiments on character‑level Tiny Shakespeare show that a model trained on 256‑character contexts achieves lower perplexity at 4096 characters using this scheme, and a second diagnostic demonstrates retrieval of multiple serialized facts.

By Mahesh Godavarti