arXiv Machine Learning By Federico Milanesio, Alessandro Ingrosso, Matteo Osella

An Analytical Theory of Auxiliary Learning

Read the original on arXiv Machine Learning →

The paper presents an analytical theory of auxiliary learning, an optimization paradigm where a neural network’s performance on a target task is enhanced by jointly training on additional tasks. Using a teacher‑student framework, the authors derive a closed system of differential equations that describe online stochastic gradient descent dynamics in the large‑input limit. For linear networks, they provide a closed‑form expression for the generalization error that shows how task correlations and label noise influence the benefit of auxiliary learning, while for nonlinear activations they develop a fluctuation‑dissipation theory linking main, auxiliary, and single‑task errors. Numerical experiments confirm the theory and illustrate how auxiliary tasks improve generalization by balancing forcing dynamics toward the optimal solution with gradient noise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

The Dynamics of Generalization in Deep Learning

arXiv:2504.16450v4 Announce Type: replace Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. T...

By Rubing Yang, Pratik Chaudhari