arXiv AI By Ekaterina Alimaskina, Gleb Molodtsov, Aleksandr Beznosikov

Analyzing Stream Collapse in Hyper-Connections: From Diagnosis to Mitigation

Read the original on arXiv AI →

arXiv:2606. 03483v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single Transformer residual stream with multiple streams, introducing a permutation symmetry over stream indices.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

How Does mHC Use Its Residual Streams? Selective Routing and Near-Identity Mixing

The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.

By Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong
arXiv Computation and Language
Aug 25

The Communication Map of a Transformer

The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.

By Richard Zhe Wang
arXiv Machine Learning
Jun 17

Dissociating Decodability and Causal Use in Bracket-Sequence Transformers

arXiv:2604. 22128v2 Announce Type: replace-cross Abstract: When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering.

By Aryan Sharma, Cutter Dawes, Shivam Raval
arXiv Machine Learning
Sep 10

LLM Layers Immediately Correct Each Other

arXiv:2609.07876v1 Announce Type: cross Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...

By Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt
arXiv Machine Learning
5d ago

The Residual Stream's Effective Depth

The paper introduces effective depth (Deff), a scalar diagnostic that treats a transformer’s layer‑wise residual stream as a discrete‑time process and measures how representation similarity decays with layer distance. Across sixteen decoder‑only language models, Deff reveals that most models exhibit a lower similarity decay than the closed‑form reference, indicating correlated residual updates rather than unused depth. The study also shows that this effect is robust to various controls and persists early in training, suggesting Deff is a global accumulated‑state diagnostic rather than a capability score.

By Barak Gahtan, Ido Galil, Alex M. Bronstein