arXiv:2608. 13510v1 Announce Type: cross Abstract: Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency.
By Nestor R. Barraza, Gabriel Pena
Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds.
arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.
By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv:2609.36271v1 Announce Type: new
Abstract: Recent work has established that power-law spectral conditions on data enable tight convergence bounds for deterministic gradient descent, resolving th...
By Thomas Dybdahl Ahle, Yaroslav Bulatov, Christopher De Sa, Christopher R\'e
arXiv:2606. 07120v1 Announce Type: new Abstract: Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error.
By Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee
arXiv:2610.00446v1 Announce Type: cross
Abstract: As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its v...
By Akshay Balsubramani
arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.
By Ori Shem-Ur, Yaron Oz
arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang
The paper presents an analytical theory of auxiliary learning, an optimization paradigm where a neural network’s performance on a target task is enhanced by jointly training on additional tasks. Using a teacher‑student framework, the authors derive a closed system of differential equations that describe online stochastic gradient descent dynamics in the large‑input limit. For linear networks, they provide a closed‑form expression for the generalization error that shows how task correlations and label noise influence the benefit of auxiliary learning, while for nonlinear activations they develop a fluctuation‑dissipation theory linking main, auxiliary, and single‑task errors. Numerical experiments confirm the theory and illustrate how auxiliary tasks improve generalization by balancing forcing dynamics toward the optimal solution with gradient noise.
By Federico Milanesio, Alessandro Ingrosso, Matteo Osella
arXiv:2606. 28879v1 Announce Type: new Abstract: The adaptive moment estimation algorithm, known as Adam, is widely used in modern machine learning, owing to its low per-iteration complexity and strong empirical performance.
By Xin Zheng, Yifei Jin, Lei Guo
arXiv:2606. 22406v2 Announce Type: replace Abstract: Attention mechanisms have demonstrated remarkable empirical success in identifying relevant information from large collections of tokens, yet the theoretical principles underlying this behavior remain poorly understood.
By Lan V. Truong
arXiv:2607. 00207v1 Announce Type: cross Abstract: We develop a framework for analyzing the learning dynamics of $\ell_2$-adversarial training of single-index models on Gaussian mixtures in the high-dimensional limit under streaming stochastic gradient descent (SGD).
By Fabrizzio Sabelli