The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates.
whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."
By Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
arXiv:2609.39408v1 Announce Type: cross
Abstract: Population loss can remain nearly constant while a neural network learns a substantially more predictive representation. We establish this separation...
By Akash Kumar
The paper presents a formula for the population loss in shallow ReLU networks with bias within the student‑teacher kernel model, extending earlier results by Choo and Saul (2009) and Brutzkus and Globerson (2017). It utilizes Owen’s T‑function, providing the necessary theory and a high‑precision implementation via MPFR. The study shows that adding bias strictly decreases loss, extends known families of spurious minima to biased networks, and indicates that the resulting change in landscape geometry is relatively mild, focusing on cases where the number of inputs equals the number of neurons.
By Michael Field
arXiv:2505. 24849v2 Announce Type: replace-cross Abstract: For three decades statistical mechanics has been providing a framework to analyse neural networks.
By Jean Barbier, Francesco Camilli, Minh-Toan Nguyen, Mauro Pastore, Rudy Skerk
arXiv:2606. 30444v1 Announce Type: cross Abstract: Neural networks are known to be susceptible to over-reliance on spurious correlations.
By Tyler LaBonte, Vidya Muthukumar
Neural networks are known to be susceptible to over-reliance on spurious correlations. However, the precise mechanism by which models exploit shortcut features is not fully understood, and algorithms to mitigate this behavior rely on as yet unjustified assumptions about the learned representations.
The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.
By Hao Yu
arXiv:2606. 05863v1 Announce Type: new Abstract: Grokking suggests that fitting the training data and learning a simple underlying rule may occur on different time scales.
By Hu Tan, Kuo Gai, Shihua Zhang
arXiv:2608. 06766v1 Announce Type: cross Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units.
By Tongxi Wang
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2601. 16884v3 Announce Type: replace Abstract: We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks.
By Shijun Zhang, Zuowei Shen, Yuesheng Xu