Full-bandwidth transformer
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.
arXiv:2607. 01218v1 Announce Type: cross Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions.
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.
The paper introduces Sparse Token Routing in Efficient Transformers, evaluating a two-stream Transformer (SEWN) that routes tokens through either lightweight or full-capacity processing via a learned gate. Experiments show that routing causes negligible accuracy change compared to parameter-matched baselines, and that the effectiveness of the gate’s token-importance signal depends on its learning method. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without altering task accuracy.
The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.
The paper introduces Prediction of Prediction (PoP), a method that fuses intermediate hidden representations across transformer layers during a single forward pass to detect hallucinations in large language models. PoP leverages internal hidden‑state transition dynamics to signal factual errors without extra decoding steps, achieving a 75.5% AUROC on the TruthfulQA benchmark with less than 1.2% added latency.
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
arXiv:2609.38149v1 Announce Type: new Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for informa...
arXiv:2605.26797v2 Announce Type: replace Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden...
arXiv:2607. 02964v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does.
arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
The paper introduces CS-MoE, a Transformer architecture that shares experts across layers to reduce inter‑layer parameter redundancy. By combining layer‑independent experts with a globally shared expert pool, CS‑MoE allows elastic control over token‑level parameter activation and computational cost. Experiments show that CS‑MoE achieves lower perplexity than equal‑scale dense Transformers while activating only 55% of parameters, and its performance scales with the number of activated experts, approaching MoE performance within a fixed FLOPs budget.
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.