arXiv:2608. 06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention.
By Bj\"orn Ladewig, Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
By Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
arXiv:2607. 15702v2 Announce Type: replace-cross Abstract: We develop a non-asymptotic approximation, sampling, and finite-iteration optimization theory for variational physics-informed approximation of uniformly monotone nonlinear multiscale elliptic equations.
By Ronald Katende
arXiv:2606. 05326v1 Announce Type: cross Abstract: We study the dynamics of gradient descent in the Edge of Stability regime, where the learning rate is large enough to induce persistent oscillations in the loss and the sharpness.
By Antonin Chodron de Courcel
arXiv:2607. 15702v1 Announce Type: cross Abstract: We prove a finite-sample formulation gap for physics-informed learning of nonlinear multiscale elliptic equations.
By Ronald Katende
arXiv:2512.12767v2 Announce Type: replace-cross
Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
By Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi
arXiv:2606. 30384v1 Announce Type: new Abstract: Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape.
By Pedro Jim\'enez-Gonz\'alez, Miguel C. Soriano, Lucas Lacasa
Training in artificial neural networks can be viewed as a trajectory evolving through a high-dimensional loss landscape. However, the large number of trainable parameters makes the direct analysis of these dynamics challenging.
arXiv:2607. 03347v1 Announce Type: new Abstract: We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}.
By Joan Bruna
arXiv:2606. 18303v1 Announce Type: cross Abstract: We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics.
By Taiki Miyagawa
Tsallis statistics generalizes Boltzmann-Gibbs statistical mechanics through a single real parameter $q$ that controls the weight assigned to rare and frequent events. Originally proposed to describe physical systems with long-range correlations, multifractal geometry, and heavy-tailed fluctuations, the framework has become a recurring ingredient in modern artificial intelligence (AI): it underlies sparse attention mechanisms (\textsc{sparsemax} and $α$-\textsc{entmax}), maximum-entropy reinforcement learning with controllable exploration, robust and heavy-tailed probabilistic models, and a family of generalized loss functions and regularizers.
arXiv:2602. 05600v2 Announce Type: replace Abstract: Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima.
By Yikuan Zhang, Ning Yang, Yuhai Tu