Scalar Representations of Neural Network Training Dynamics
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
arXiv:2602. 14885v2 Announce Type: replace-cross Abstract: Recurrent neural networks (RNNs) provide a theoretical framework for understanding computation in biological neural circuits, yet classical results, such as Hopfield's model of associative memory, rely on symmetric connectivity that restricts network dynamics to gradient-like flows.
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging.
arXiv:2507.10383v5 Announce Type: replace-cross Abstract: Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neu...
arXiv:2507. 05164v2 Announce Type: replace-cross Abstract: In this chapter, we utilize dynamical systems to analyze several aspects of machine learning algorithms.
The book "The Principles of Diffusion Models" outlines the foundational concepts behind diffusion models, tracing their evolution from a forward process that corrupts data into noise to a reverse process that reconstructs data. It presents three complementary perspectives—variational, score-based, and flow-based—each describing how a time-dependent velocity field transports a simple prior to the data distribution. The text also covers practical guidance for controllable generation, efficient solvers, and diffusion-inspired flow-map models, providing a mathematically grounded framework for readers with basic deep‑learning knowledge.
arXiv:2510. 09685v2 Announce Type: replace-cross Abstract: Deep learning has become a pivotal technology in fields such as computer vision, scientific computing, and dynamical systems, significantly advancing these disciplines.
arXiv:2606. 21295v2 Announce Type: replace-cross Abstract: Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wise dynamics, where all neurons in the same layer co-evolve through a shared parameterized operator, leaving individual neurons no freedom to evolve independently.
The paper presents a non-equilibrium dynamical mean-field theory (DMFT) that explains how learning reshapes the dynamics of recurrent neural networks, turning initially chaotic activity into stable, task-dependent behavior. It shows that a slow, feedback-driven learning process gradually increases effective feedback strength, driving the network through a bifurcation that marks the transition from chaotic to stable dynamics. By deriving the two-time correlation function, the authors identify a critical feedback strength and a learning-rate-dependent critical time that separate these regimes, and they demonstrate that the theory accurately predicts the network’s output evolution during training, matching numerical simulations.
Nonlinear GENERIC-Embedded Neural Networks (N-GENNs) are a deep learning framework designed to discover evolution equations for systems governed by the nonlinear GENERIC formalism. The method incorporates generalized gradient flows through convex dissipation potentials, allowing it to capture a wider range of thermodynamically consistent dynamics, including those with non‑quadratic dissipation potentials. Thermodynamic structure is enforced by construction, ensuring compliance with the first and second laws, and the approach is validated on a harmonic oscillator with a heat bath, an idealized chemical motor, and a one‑dimensional viscoplastic Perzyna model.
arXiv:2606. 10530v1 Announce Type: cross Abstract: Recent developments in brain recording are driving a demand for machine learning tools capable of decoding the latent structure of large populations of neurons.
arXiv:2606. 05272v1 Announce Type: new Abstract: Neural rough differential equations (NRDEs) stay accurate under irregular sampling while taking far fewer integration steps than standard neural differential equations, summarising a finely sampled driver by its log-signature and advancing the hidden state over coarse intervals using the log-ODE method.
arXiv:2602. 15649v2 Announce Type: replace Abstract: In dynamical systems reconstruction (DSR) we aim to recover the dynamical system (DS) underlying observed time series.