Tokenization in Transformers v5: Simpler, Clearer, and More Modular
Related stories
tokenizers v1: encode, decode and scaling, measured
Distilling Sequential Computation in Transformer Language Models
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits
arXiv:2607. 07066v1 Announce Type: cross Abstract: Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition.
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
arXiv:2606. 20076v1 Announce Type: cross Abstract: Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the tokenizer's fixed compression ratio.
Distilling Sequential Computation in Transformer Language Models
arXiv:2609.27233v1 Announce Type: new Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...
Sparse Token Routing in Efficient Transformers
The paper introduces Sparse Token Routing in Efficient Transformers, evaluating a two-stream Transformer (SEWN) that routes tokens through either lightweight or full-capacity processing via a learned gate. Experiments show that routing causes negligible accuracy change compared to parameter-matched baselines, and that the effectiveness of the gate’s token-importance signal depends on its learning method. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without altering task accuracy.
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
arXiv:2511. 17864v3 Announce Type: replace Abstract: Recent research has established that the impact of context in a vanilla transformer can be represented implicitly by forming a token-dependent, rank-1 patch to its MLP weights.
Fixed Universal Transformers
The paper introduces fixed universal transformers, which are transformers with immutable internal parameters that can emulate any transformer within a specified class by encoding the target model’s description into the input embedding. The authors provide explicit sparse constructions that achieve universality when the embedding dimension is large enough, and demonstrate that universality is generic—randomly initialized transformers are almost surely universal. Empirical tests on parenthesis balancing and multi‑hop reasoning tasks support the theory, suggesting that a transformer’s expressive power largely stems from its input representation rather than its learned weights.
On the Expressive Power of Transformers
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers
The paper investigates how the choice of graph tokenization affects transformer expressivity. It analyzes three tokenization families—spectral, random‑walk, and adjacency—showing that each induces different depth requirements and that some tokenizations are inherently lossy or ill‑conditioned for certain tasks. The authors prove lower bounds and impossibility results for converting between tokenizations and validate these findings with experiments on synthetic and real‑world data.
Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions
arXiv:2606. 08768v1 Announce Type: new Abstract: Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings.