arXiv Machine Learning

Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation

arXiv:2501. 18530v3 Announce Type: replace-cross Abstract: We consider a teacher-student model of supervised learning with a fully-trained two-layer neural network whose width $k$ and input dimension $d$ are large and proportional.

arXiv Machine Learning
Jul 30

On the robustness of noisy solutions in non-convex neural networks

arXiv:2607. 27000v1 Announce Type: cross Abstract: Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare.

By Enrico M. Malatesta, Alessandra Passalacqua, Riccardo Zecchina