arXiv Machine Learning

Population loss in shallow ReLU networks: Bias & families of critical points

The paper presents a formula for the population loss in shallow ReLU networks with bias within the student‑teacher kernel model, extending earlier results by Choo and Saul (2009) and Brutzkus and Globerson (2017). It utilizes Owen’s T‑function, providing the necessary theory and a high‑precision implementation via MPFR. The study shows that adding bias strictly decreases loss, extends known families of spurious minima to biased networks, and indicates that the resulting change in landscape geometry is relatively mild, focusing on cases where the number of inputs equals the number of neurons.

arXiv AI
Sep 11

Teacher Geometry Shapes Learnability in Teacher-Student Networks

The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates. whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."

By Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea
arXiv Machine Learning
Jun 16

Constraining the outputs of ReLU neural networks

arXiv:2508. 03867v2 Announce Type: replace-cross Abstract: We introduce a class of algebraic varieties naturally associated with ReLU neural networks, arising from the piecewise linear structure of their outputs across activation regions in input space, and the piecewise multilinear structure in parameter space.

By Yulia Alexandr, Guido Mont\'ufar
arXiv Machine Learning
Jul 14

Approximation of Analytic Functions by ReLU Neural Networks with Adjustable Depth and Width

arXiv:2607. 10589v1 Announce Type: cross Abstract: In contrast to most studies on neural network approximation theory that characterize results through a single parameter, such as the total number of network parameters, \cite{shen2020deep} pioneered the characterization of approximation rates as a joint function of the width parameter $N$ and the depth parameter $L$, thereby granting greater architectural flexibility.

By Yanming Lai, Defeng Sun, Yang Wang