arXiv Machine Learning

xHC: Expanded Hyper-Connections

arXiv:2607. 14530v1 Announce Type: new Abstract: Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth.

arXiv AI
Sep 7

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.

By Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong
arXiv Machine Learning
Sep 3

oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions

arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.

By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo
arXiv AI
Aug 25

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

The paper proposes a theoretical framework for scheduling high‑quality data in large language model training by extending functional scaling laws to account for time‑varying data quality. It identifies two regimes—noise‑limited and signal‑limited—where high‑quality data should be used differently, and introduces a Drop‑Stable‑Rampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.

By Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu
arXiv Machine Learning
Jun 10

PRISM: Parallel Residual Iterative Sequence Model

arXiv:2602. 10796v3 Announce Type: replace Abstract: Generative sequence modeling faces a fundamental tension between the expressivity of Transformers and the efficiency of linear sequence models.

By Jie Jiang, Ke Cheng, Xin Xu, Mengyang Pang, Tianhao Lu, Jiaheng Li, Yue Liu, Yuan Wang, Jun Zhang, Huan Yu, Zhouchen Lin
Hugging Face Trending Papers
Sep 2

oHC: Orthogonal Hyper-Connections on SO(4) via Quaternions

The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.