arXiv:2610.01728v1 Announce Type: cross
Abstract: Understanding loss landscapes is central to explaining neural-network training, yet their structure remains only partially understood even in simple...
By Jakob Paul Zimmermann, Moritz Grillo, Andrei Balakin, Georg Loho
The paper investigates how the geometry of teacher neural networks affects the learnability of student networks in teacher‑student setups. By formalizing learnability as the success rate of reaching the global minimum, the authors identify two teacher distributions—one maximizing node dissimilarity (easy) and one minimizing it (hard)—that lead to markedly different success rates across various settings and activation functions. They analyze the loss landscape of small networks, revealing two types of suboptimal local minima (out‑of‑bounds and interior) whose attraction regions depend on teacher structure, and demonstrate that adjusting learning rates for the readout layer and inner biases can improve success rates.
whyItMatters:"The study highlights that teacher geometry, often overlooked, plays a crucial role in determining how effectively a student network can learn, offering guidance for designing more realistic teacher‑student experiments."
By Kai J. Sandbrink, Flavio Martinelli, Alexander van Meegen, Wulfram Gerstner, Johanni Brea
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2606. 07728v1 Announce Type: new Abstract: It is well established that ReLU networks define continuous piecewise-linear functions, and that their linear regions are polyhedra in the input space.
By Blake B. Gaines, Jinbo Bi
arXiv:2410. 00722v3 Announce Type: replace Abstract: We study convolutional neural networks with monomial activation functions.
By Vahid Shahverdi, Giovanni Luca Marchetti, Kathl\'en Kohn
arXiv:2606. 04327v1 Announce Type: cross Abstract: We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions.
By Tian Ding, Dawei Li, Ruoyu Sun