Scaling depth capacity via zero/one-layer model expansion
arXiv:2511. 04981v2 Announce Type: replace Abstract: Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost.
arXiv:2607. 14530v1 Announce Type: new Abstract: Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth.
arXiv:2511. 04981v2 Announce Type: replace Abstract: Model depth is a double-edged sword in deep learning: deeper models achieve higher accuracy but require higher computational cost.
The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.
arXiv:2606. 16112v1 Announce Type: cross Abstract: Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual stream can grow rapidly with depth.
arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed.
arXiv:2606. 01117v1 Announce Type: cross Abstract: Extreme multi-label classification (XMC) involves learning models over large output spaces with millions of labels, making the output layer a memory-compute bottleneck.
arXiv:2605. 17842v3 Announce Type: replace Abstract: Autoregressive language models execute Transformer layers sequentially, creating a latency bottleneck that is not removed by conventional tensor or pipeline parallelism.
The paper proposes a theoretical framework for scheduling high‑quality data in large language model training by extending functional scaling laws to account for time‑varying data quality. It identifies two regimes—noise‑limited and signal‑limited—where high‑quality data should be used differently, and introduces a Drop‑Stable‑Rampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.
arXiv:2602. 10796v3 Announce Type: replace Abstract: Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models.
arXiv:2606. 26744v1 Announce Type: new Abstract: We present HyperDFlash, a block-parallel speculative decoding framework tailored to the novel multi-hyper-connection (MHC) architecture proposed by DeepSeek-V4.
arXiv:2606. 26538v1 Announce Type: cross Abstract: Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value.
The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.