We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.
arXiv:2607. 14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization.
By Katie Everett
arXiv:2609.37680v1 Announce Type: cross
Abstract: One of the current premises of mechanistic interpretability research is that detailed accounts of the geometry of neural network representations can...
By Sai Sumedh R. Hindupur, Hadas Orgad, Thomas Fel, Demba Ba
arXiv:2604. 14037v2 Announce Type: replace Abstract: Parameter space is not function space for neural network architectures.
By Pranavkrishnan Ramakrishnan
arXiv:2608. 02816v1 Announce Type: new Abstract: We study the topology of learned representations in predictive coding networks (PCNs), a neuro-inspired bidirectional architecture, using a quantitative layer-wise persistent homology analysis.
By Adam Shaw, Jiayu Li, Michael Sperling, Michael Kim, Alvin Jin
arXiv:2609.39078v1 Announce Type: new
Abstract: Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and ar...
By Marvin Theiss, Lukas Braun, Andrew M. Saxe, Erin Grant