arXiv:2606. 03483v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single Transformer residual stream with multiple streams, introducing a permutation symmetry over stream indices.
By Ekaterina Alimaskina, Gleb Molodtsov, Aleksandr Beznosikov
arXiv:2607. 14530v1 Announce Type: new Abstract: Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth.
By Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan
The paper introduces OASIS, a method designed to stabilize dual‑normalized attention‑residual architectures by employing explicit null routing and token‑to‑depth null coupling. OASIS mitigates attention sinks and activation outliers, improving low‑bit quantization performance across several language‑model backbones. Empirical results show significant reductions in attention norms and perplexity, with notable gains on long‑context benchmarks.
By Haozheng Luo, Haoran Dai, Jingyuan Huang, Shaoyang Zhang, Xi Chen, Eric Hanchen Jiang, Yijiang Li, Chenghao Qiu, Chenwei Xu, Zhenyu Pan, Haotian Zhang, Binghui Wang, Yan Chen
arXiv:2606. 26538v1 Announce Type: cross Abstract: Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value.
By Huzama Ahmad, Cao Viet Hai Nam, Se-Young Yun
arXiv:2606. 06564v1 Announce Type: cross Abstract: Residual connections are central to training deep Transformers, but standard PreNorm residual streams aggregate sublayer updates with fixed unit weights.
By Kehan Wang
arXiv:2606. 26744v1 Announce Type: new Abstract: We present HyperDFlash, a block-parallel speculative decoding framework tailored to the novel multi-hyper-connection (MHC) architecture proposed by DeepSeek-V4.
By Luxi Lin, Shuang Peng, Rui Ma, Junhao Hua, Shuwei Fan, Zhengda Qin, Qiang Wang, Hongjian Sun, Fangmin Chen, Songwei Liu
The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.
By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo
The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.
By Richard Zhe Wang
arXiv:2606. 00091v1 Announce Type: cross Abstract: Joint Embedding Predictive Architectures (JEPAs) have reshaped self-supervised representation learning in vision.
By Sangdae Nam
arXiv:2608. 15787v1 Announce Type: cross Abstract: Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts.
By Cedric Caruzzo, Donggeun Yoo, Tae Soo Kim
arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
By Nelson Guda