SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
arXiv:2511. 17864v3 Announce Type: replace Abstract: Recent research has established that the impact of context in a vanilla transformer can be represented implicitly by forming a token-dependent, rank-1 patch to its MLP weights.
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
The paper introduces CS-MoE, a Transformer architecture that shares experts across layers to reduce inter‑layer parameter redundancy. By combining layer‑independent experts with a globally shared expert pool, CS‑MoE allows elastic control over token‑level parameter activation and computational cost. Experiments show that CS‑MoE achieves lower perplexity than equal‑scale dense Transformers while activating only 55% of parameters, and its performance scales with the number of activated experts, approaching MoE performance within a fixed FLOPs budget.
arXiv:2604. 21254v3 Announce Type: replace Abstract: LLM architecture research generally aims to maximize model quality subject to fixed compute/latency budgets.
arXiv:2610. 01172v1 Announce Type: new Abstract: We study learning rate (LR) scaling for hybrid architectures combining Transformer and State-Space Model (SSM) blocks, a class adopted by several recent production language models.
arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.
arXiv:2511. 12081v2 Announce Type: replace-cross Abstract: Despite massive investments in scale, deep models for click-through rate (CTR) prediction often exhibit rapidly diminishing returns -- a stark contrast to the {predictable scaling laws} seen in large language models (LLMs).
arXiv:2606. 27449v1 Announce Type: new Abstract: Multi-head attention conventionally partitions the hidden dimension equally across all heads at every layer, enforcing an identical representational subspace dimension (dh = dmodel/h) throughout the models depth.
arXiv:2607. 22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes.
arXiv:2507. 16003v4 Announce Type: replace-cross Abstract: One of the most striking features of Large Language Models (LLMs) is their ability to learn in-context.
arXiv:2607. 22757v1 Announce Type: cross Abstract: We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grading and propagates the induced weighted scalar action through embeddings, self-attention, and the training objective.
The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.
arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.