Scalar Representations of Neural Network Training Dynamics
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
arXiv:2605. 12763v2 Announce Type: replace Abstract: Rich learning in recurrent neural networks often proceeds through sudden transitions in latent dynamics, but there is little theory predicting how gradient descent behaves during these events.
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
arXiv:2606. 15551v1 Announce Type: new Abstract: The Edge of Stability (EoS) phenomenon, where gradient descent operates with sharpness exceeding the classical convergence threshold yet the loss decreases over long timescales, is ubiquitous in modern deep learning but remains poorly understood in realistic settings.
Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging.
arXiv:2601. 19019v3 Announce Type: replace-cross Abstract: Neural population activity in sensory cortex is organized on low-dimensional manifolds, but why such manifolds arise and what determines their geometry remain unclear.
arXiv:2606. 10530v1 Announce Type: cross Abstract: Recent developments in brain recording are driving a demand for machine learning tools capable of decoding the latent structure of large populations of neurons.
arXiv:2607. 02283v1 Announce Type: cross Abstract: In-context learning (ICL) operates via implicit gradient descent embedded in the forward pass of modern AI architectures -- Transformers, Mamba, state-space models, and MLPs.
arXiv:2607. 14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization.
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.
arXiv:2606. 09929v1 Announce Type: cross Abstract: Physical reservoir computing harnesses nonlinear mechanical dynamics but, by convention, freezes the substrate and trains only a linear readout, presuming the substrate is not usefully trainable.
arXiv:2608. 14803v1 Announce Type: new Abstract: A recent line of work recasts the post-memorization phase of grokking as constrained optimization: once a network interpolates the training set, weight decay drives a slow drift along the zero-loss manifold toward lower norm.
arXiv:2607. 04993v1 Announce Type: cross Abstract: Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selects among them.
arXiv:2606. 24396v1 Announce Type: new Abstract: Large Transformer models function as Dense Associative Memories (DAMs), retrieving knowledge via high-dimensional attractor dynamics driven by the self-attention mechanism \citep{ramsauer2020hopfield, wu2024attention}.