arXiv:2604. 00316v2 Announce Type: replace-cross Abstract: Grokking occurs when a model achieves high training accuracy but generalization to unseen test points happens long after that.
By Marcel Tom\`as Bernal, Neil Rohit Mallinar, Mikhail Belkin
arXiv:2604. 15613v4 Announce Type: replace-cross Abstract: We present Green-ELM, a non-iterative neural architecture that replaces gradient-based optimization of the output layer with a closed-form analytic solution over a fixed, high-dimensional random feature representation.
By Wladimir Silva
arXiv:2608. 14733v1 Announce Type: cross Abstract: Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier features using $\cos$, $\sin$, or a combination of both.
By Qihong Yang, Zhijie Su, Yangtao Deng, Qiaolin He
arXiv:2506.11030v2 Announce Type: replace-cross
Abstract: Training neural networks has traditionally relied on backpropagation (BP), a gradient-based algorithm that, despite its widespread success, s...
By Nazmus Saadat As-Saquib, A N M Nafiz Abeer, Hung-Ta Chien, Byung-Jun Yoon, Suhas Kumar, Su-in Yi
arXiv:2311. 02960v5 Announce Type: replace Abstract: Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data.
By Peng Wang, Xiao Li, Can Yaras, Zhihui Zhu, Laura Balzano, Wei Hu, Qing Qu
arXiv:2606. 11319v1 Announce Type: new Abstract: Learning from imperfect data is a central theme in machine learning, connecting practical questions of robustness to fundamental questions of learnability.
By Justin Tahmassebpur, Asadullah Bhuiyan, Hyejin Kim, Omri Lesser
arXiv:2608. 07043v1 Announce Type: cross Abstract: Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identification.
By Albert Saiapin, Kim Batselier
The paper investigates how multilayer perceptrons (MLPs) learn features in regression tasks with clustered data. It finds that instead of forming a single global low‑dimensional representation, MLPs develop monosemantic specialized neurons—each neuron aligns strongly with a specific predictive feature relevant to a particular region of the input space. This specialization results in a collection of local low‑dimensional representations, giving MLPs a provable data‑efficiency advantage over methods that rely on a global representation.
By Amirhesam Abedsoltan, Enric Boix-Adsera, Fivos Kalogiannis, Mikhail Belkin
arXiv:2506.07188v2 Announce Type: replace
Abstract: End-to-end neural networks have become a dominant paradigm in autonomous driving, where reliable deployment requires controllable post-training ada...
By Ni Ding, Shuchang Wang, Lei He, Shengbo Eben Li, Keqiang Li
Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?
arXiv:2502. 07209v4 Announce Type: replace Abstract: Physics-Informed Neural Networks (PINNs) seek to solve partial differential equations (PDEs) with deep learning.
By Shaghayegh Fazliani, Zachary Frangella, Madeleine Udell
arXiv:2605. 06938v2 Announce Type: replace-cross Abstract: Recently Brown et al.
By Brian Charles Brown, Mauricio Munoz, Robert Bridges, David Grimsman, Sean Warnick