Selective Rotary Position Embedding
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
arXiv:2509. 10534v3 Announce Type: replace-cross Abstract: The attention mechanism in a Transformer architecture matches key to query based on both content -- the what -- and position in a sequence -- the where.
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
The paper proposes an encoder Transformer that explicitly separates semantic, absolute positional (AP), and relative positional (RP) information, restricting the masked‑language‑modeling objective to the semantic stream. This disentanglement reveals that the AP subspace collapses into a low‑frequency two‑dimensional manifold reflecting document structure, that attention heads specialize into structure‑ and semantic‑oriented groups with RP supporting only the latter, and that standard positional encodings fail to robustly encode macroscopic structure. The approach preserves positional encoding and improves performance on 49 out of 65 linguistic phenomena in the Flash‑Holmes probing benchmark.
arXiv:2609.38109v1 Announce Type: cross Abstract: The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of...
The paper proposes a content‑based addressing scheme for long‑context models that replaces the growing token counter in Rotary Position Embedding (RoPE) with unit‑level addresses derived from the content of each unit. By dividing the token stream into units, the method preserves local RoPE behavior while allowing new units to be addressed via learned content maps, avoiding positional mismatches when extending context length. Experiments on character‑level Tiny Shakespeare show that a model trained on 256‑character contexts achieves lower perplexity at 4096 characters using this scheme, and a second diagnostic demonstrates retrieval of multiple serialized facts.
arXiv:2607. 19363v1 Announce Type: new Abstract: Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a uniform frequency schedule and scaling across all attention heads.
arXiv:2607. 10134v1 Announce Type: new Abstract: Rotary Positional Encodings (RoPE) are currently the most popular positional encodings used in modern language models.
T‑RoPE introduces a time‑aware Rotary Position Embedding for sequential recommendation, replacing index‑only rotations with timestamp‑based angles, learnable temporal coefficients, multiscale frequency banks, shifted query alignment, and non‑stationary key rotation. The authors prove that standard RoPE is time‑translation invariant and cannot capture seasonal contexts, while T‑RoPE breaks this invariance while preserving the RoPE interface. Across five public benchmarks and an industrial‑scale e‑commerce dataset, T‑RoPE outperforms baselines by up to 130 % in HR@10 and delivers significant online lift in conversion and order count.
The paper introduces DEPT, a method that trains a single decoder-only large language model to both expand queries and encode documents for retrieval. By preserving document embeddings close to their initial cached values while allowing gradients to flow through the generator, DEPT stabilizes retrieval targets and enables efficient index reuse and online hard‑negative mining. Experiments on the BEIR benchmark with Qwen3‑4B‑Instruct‑2507 and LLaMA‑3.2‑3B‑Instruct show that DEPT outperforms training‑free, independently trained, and staged unified baselines, with ablations confirming the benefits of preservation, whitening, end‑to‑end expansion training, and online negatives.
arXiv:2607. 07678v1 Announce Type: new Abstract: Rotary Position Embeddings (RoPE) provide transformers with a fixed grid of positional frequencies, yet trained models use these frequencies highly non-uniformly.
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
The paper proposes a new architecture for masked language modeling that replaces the Transformer attention mechanism with a stack of low‑rank bottleneck autoencoders. Each autoencoder mixes information locally, across the full sequence, and across attention heads, compressing and reconstructing inputs without training‑dependent width. An iterative refinement process at masked positions pulls embeddings toward a weighted neighbor average and then projects them back onto the learned manifold, achieving comparable performance to BERT with roughly 1.9× fewer FLOPs and matching BERT on rare‑token performance through a frequency‑aware training schedule.
arXiv:2609.39929v1 Announce Type: cross Abstract: Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and disting...