arXiv:2601. 21579v2 Announce Type: replace-cross Abstract: The success of Hyper-Connections (HC) in neural networks (NN) has also highlighted issues related to training instability and restricted scalability.
By Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Danilo Mandic
arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.
By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo
The paper introduces Orthogonal Hyper-Connections (oHC), a new approach that replaces the single residual stream of a Transformer with multiple parallel streams mixed by a rotation matrix from the group SO(n). By constraining the mixing matrix to SO(n) and parameterizing it with unit quaternions for four streams, oHC prevents both amplification and attenuation of residuals, maintaining training stability and preserving stream diversity. Experiments show that oHC outperforms the single-stream baseline, manifold-constrained Hyper-Connections, and identity-fixed Hyper-Connections across a wide range of downstream tasks.
arXiv:2608. 07851v1 Announce Type: new Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks.
By Yuxuan Gu, Wuyang Zhou, Huijun Xing, Danilo Mandic
arXiv:2607. 21885v1 Announce Type: new Abstract: Coarsening-based training for graph neural networks (GNNs), i.
By Guoming Li, Jian Yang, Xukun Wang, Zixiao Wang, Shangsong Liang, Yifan Chen
The paper introduces Spectral‑Sphere‑Constrained Hyper‑Connections (s²HC), a new method for controlling the residual matrices used in Hyper‑Connections (HC). Unlike previous doubly stochastic constraints that caused identity degeneration, expressivity bottlenecks, and parameterization inefficiencies, s²HC confines these matrices to a spectral norm sphere, restoring flexibility over subdominant spectra and eliminating unstable Sinkhorn‑Knopp iterations. This approach preserves training stability while allowing expressive, non‑degenerate residual matrices.
By Zhaoyi Liu, Haichuan Zhang, Ang Li