Can Representation Learning Decouple from Loss Minimization? Polar Updates Have an Answer
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
Direct feedback alignment (DFA) trains hidden layers via fixed random projections of output error, but with tanh hidden units and independent sigmoid outputs, plain stochastic gradient descent can stall near a constant predictor of class frequencies. This stall is traced to the error’s common mode—a rank‑one component shared across inputs—that drives tanh units toward saturation. The study shows that calibration of the baseline readout to class priors suppresses collapse and speeds learning, while other interventions such as using Adam, adjusting feedback strength, or subtracting batch means affect the severity and recovery of collapse across MNIST, CIFAR‑10, and deeper networks.
arXiv:2606. 05675v1 Announce Type: new Abstract: Continual learning (CL) seeks models that acquire new skills without erasing prior knowledge.
arXiv:2606. 25971v1 Announce Type: new Abstract: Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object.
arXiv:2608. 05136v1 Announce Type: new Abstract: Gradient descent on a factored model $W = UV^\top$ is implicitly biased toward low-rank solutions, while Adam, starting from the same small initialization, is not.
arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.
arXiv:2609.08381v1 Announce Type: cross Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...