arXiv Machine Learning

Speed Limit for Information Acquisition in Stochastic Learning Dynamics

Hugging Face Trending Papers
Aug 13

On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective

Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds.

arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv Statistics ML
2d ago

Exact information accounting for SGD methods

arXiv:2610.00446v1 Announce Type: cross Abstract: As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its v...

By Akshay Balsubramani
arXiv Machine Learning
Sep 25

An Analytical Theory of Auxiliary Learning

The paper presents an analytical theory of auxiliary learning, an optimization paradigm where a neural network’s performance on a target task is enhanced by jointly training on additional tasks. Using a teacher‑student framework, the authors derive a closed system of differential equations that describe online stochastic gradient descent dynamics in the large‑input limit. For linear networks, they provide a closed‑form expression for the generalization error that shows how task correlations and label noise influence the benefit of auxiliary learning, while for nonlinear activations they develop a fluctuation‑dissipation theory linking main, auxiliary, and single‑task errors. Numerical experiments confirm the theory and illustrate how auxiliary tasks improve generalization by balancing forcing dynamics toward the optimal solution with gradient noise.

By Federico Milanesio, Alessandro Ingrosso, Matteo Osella