arXiv AI

Analyzing Stream Collapse in Hyper-Connections: From Diagnosis to Mitigation

arXiv:2606. 03483v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single Transformer residual stream with multiple streams, introducing a permutation symmetry over stream indices.

arXiv AI
Sep 7

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.

By Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong
arXiv Computation and Language
Aug 25

The Communication Map of a Transformer

The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.

By Richard Zhe Wang
arXiv Machine Learning
Jun 17

Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

arXiv:2604. 22128v2 Announce Type: replace-cross Abstract: When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering.

By Aryan Sharma, Cutter Dawes, Shivam Raval
arXiv Machine Learning
Sep 10

LLM Layers Immediately Correct Each Other

arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...

By Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt
arXiv Machine Learning
5d ago

The Residual Stream's Effective Depth

The paper introduces effective depth (Deff), a scalar diagnostic that treats a transformer’s layer‑wise residual stream as a discrete‑time process and measures how representation similarity decays with layer distance. Across sixteen decoder‑only language models, Deff reveals that most models exhibit a lower similarity decay than the closed‑form reference, indicating correlated residual updates rather than unused depth. The study also shows that this effect is robust to various controls and persists early in training, suggesting Deff is a global accumulated‑state diagnostic rather than a capability score.

By Barak Gahtan, Ido Galil, Alex M. Bronstein
arXiv AI
Sep 18

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

The paper critiques the common practice of evaluating depth usage in depth‑recurrent language models by truncating depth during inference and measuring performance decline. It argues that this method conflates three distinct effects—fewer block applications, reduced computation, and an out‑of‑distribution readout—yet is usually interpreted as measuring only the second. To address this, the authors introduce the Depth Control Protocol (DCP), a suite of positive and negative controls that isolate each factor, along with a training intervention to confirm causality, specifically tailored for depth‑wise weight‑sharing architectures.

By Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung