arXiv Machine Learning By Xingjian Wang, Qingyu Han, Xiaodong Luo, Yin Zhang

A Mechanistic Diagnostic of Rank Collapse in Post-Norm Decoder Transformers

Read the original on arXiv Machine Learning →

arXiv:2608. 09417v1 Announce Type: new Abstract: Deep decoder-only Transformers often replace the original Post-Norm architecture with Pre-Norm variants because Post-Norm training is highly sensitive to warmup and learning rate under conventional initialization schemes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 15

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.

arXiv Machine Learning
5d ago

Common-Mode Collapse and Recovery in Direct Feedback Alignment

Direct feedback alignment (DFA) trains hidden layers via fixed random projections of output error, but with tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near a constant predictor of class frequencies. This stall is traced to the error’s common mode—a rank‑one component shared across inputs—that drives tanh units toward saturation. The study shows that calibration of the baseline readout to class priors suppresses collapse and speeds learning, while other interventions such as using Adam, adjusting feedback strength, or subtracting batch means affect the severity and recovery of collapse across MNIST, CIFAR‑10, and deeper networks.

By Varun Reddy, Bernardo L. Sabatini, Houman Safaai