The paper investigates how the four‑stream manifold‑constrained hyper‑connection (mHC) residual pathway in DeepSeek‑V4‑Flash is actually used. It finds that read/write routing is concentrated, typically involving only two streams per block, and that the dominant stream shifts across layers while representations stay directionally distinct. Residual mixing is modest, mainly in early layers, and late mixing contributes little to performance, whereas early mixing is crucial for perplexity and task scores.
By Pengxiang Zhao, Xing Li, Xianzhi Yu, Wei Guo, Zhenhua Dong
The paper introduces the Communication Map, a method that charts every potential communication channel in a transformer model using only its weights. It generalizes previous coupling metrics into a single coefficient covering all 18 connection classes, revealing that 70‑89% of head pairs are non‑randomly oriented and identifying strong or avoiding couplings. The authors demonstrate the map’s utility by recovering known induction circuits and uncovering a two‑dimensional stream subspace whose removal eliminates induction capabilities across several models.
By Richard Zhe Wang
arXiv:2604. 22128v2 Announce Type: replace-cross Abstract: When trained on tasks requiring an understanding of hierarchical structure, transformers have been found to represent this hierarchy in distinct ways: in the geometry of the residual stream, and in stack-like attention patterns maintaining a last-in, first-out ordering.
By Aryan Sharma, Cutter Dawes, Shivam Raval
arXiv:2609.07876v1 Announce Type: cross
Abstract: Recent methods in language model interpretability employ techniques such as sparse autoencoders to decompose residual stream contributions into linea...
By Arjun Patrawala, Jiahai Feng, Erik Jones, Jacob Steinhardt
arXiv:2609. 12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities.
By Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte, Holger Giese
The paper introduces effective depth (Deff), a scalar diagnostic that treats a transformer’s layer‑wise residual stream as a discrete‑time process and measures how representation similarity decays with layer distance. Across sixteen decoder‑only language models, Deff reveals that most models exhibit a lower similarity decay than the closed‑form reference, indicating correlated residual updates rather than unused depth. The study also shows that this effect is robust to various controls and persists early in training, suggesting Deff is a global accumulated‑state diagnostic rather than a capability score.
By Barak Gahtan, Ido Galil, Alex M. Bronstein
arXiv:2610.00024v1 Announce Type: new
Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer inte...
By Genpei Zhang
arXiv:2609.37824v1 Announce Type: cross
Abstract: Transformer language models build predictions through successive residual updates, but how their representations become specific to an eventual outco...
By Timur Mudarisov, Mikhail Burtsev, Radu State
arXiv:2608. 12447v1 Announce Type: new Abstract: Trained transformer models develop privileged bases: coordinate axes whose statistics differ from the rest of the residual stream.
By Nelson Guda
The paper critiques the common practice of evaluating depth usage in depth‑recurrent language models by truncating depth during inference and measuring performance decline. It argues that this method conflates three distinct effects—fewer block applications, reduced computation, and an out‑of‑distribution readout—yet is usually interpreted as measuring only the second. To address this, the authors introduce the Depth Control Protocol (DCP), a suite of positive and negative controls that isolate each factor, along with a training intervention to confirm causality, specifically tailored for depth‑wise weight‑sharing architectures.
By Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung
arXiv:2607. 20652v1 Announce Type: cross Abstract: Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual streams.
By Andrew Mack, Kraig Yuheng Tou, Mark Henry, Zhengxun Wu, Lauren Greenspan
arXiv:2609.38802v1 Announce Type: cross
Abstract: Uncontrolled repetition can prolong autoregressive generation in large language models (LLMs) and enable resource consumption attacks. Prior analyses...
By Yuanhe Zhang, Xinyao Zhou, Haoran Gao, Yuyao Zhang, Zhenhong Zhou, Fanyu Meng, Li Sun, Sen Su