Selective Rotary Position Embedding
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
arXiv:2510. 18315v2 Announce Type: replace-cross Abstract: We investigate how embedding dimension affects the emergence of an internal "world model" in a transformer trained with reinforcement learning to perform bubble-sort-style adjacent swaps.
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.
arXiv:2510. 25013v2 Announce Type: replace-cross Abstract: Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits.
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous works on Transformer learning dynamics have so far been mostly tied to specific tasks, we study a generalized class of inductive tasks that unifies several synthetic tasks known in the literature, including in-context n-grams and multi-hop reasoning.
The paper introduces ReST, a recommendation‑native Transformer scaling framework designed to handle noisy, irregular, and sparsely supervised user behavior sequences in production ranking. ReST employs a dual‑gated attention encoder with rotary positional and temporal embeddings, and a lightweight cross decoder that decouples heavy encoding from fast decoding, enabling efficient compute‑once, decode‑many‑times ranking. Experiments on industrial and public benchmarks show that ReST outperforms traditional Transformer blocks, achieving higher accuracy and consistent scaling across sequence length, depth, and width, and a one‑week online A/B test on a production advertising platform yielded a 1.31% AUC lift and an 11.93% increase in a core revenue metric within a 50 ms P99 latency budget.
arXiv:2609.38109v1 Announce Type: cross Abstract: The attention operation is naively position invariant. However, positional information is fundamental to natural language, and therefore a variety of...
arXiv:2607. 11875v1 Announce Type: cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models.
arXiv:2502.09245v3 Announce Type: replace Abstract: In contrast to RNNs, which compress their history into a single hidden state, Transformers can attend to all past tokens directly. However, standar...
arXiv:2510.19266v3 Announce Type: replace Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
arXiv:2505. 15548v2 Announce Type: replace Abstract: Autoregressive transformer language models frequently exhibit training instability when trained on long sequences, particularly under low-precision arithmetic.
arXiv:2606. 18587v1 Announce Type: cross Abstract: Decoder-only Transformers compute attention over the KV cache of preceding tokens.
arXiv:2609.07086v1 Announce Type: new Abstract: Transformers are a dominant architecture in modern machine learning, powering applications across vision, language, and beyond. At the core of their su...