arXiv:2607. 04135v1 Announce Type: cross Abstract: The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points.
By Chan Li, Nigel Goldenfeld
arXiv:2607. 05735v1 Announce Type: cross Abstract: Infinite-width limits are a standard way to reason about neural networks, but it is not automatic that the limiting learner has the same complexity-theoretic inductive bias as large finite networks.
By Dmitry Vaintrob, Kaarel H\"anni
arXiv:2607. 07127v1 Announce Type: cross Abstract: Lattice field theory is the workhorse of non-perturbative physics, used to simulate phenomena from the strong nuclear force to critical phenomena in materials.
By Tobias G\"obel, Julian R. Ebelt, Zier Mensch, Mathis Gerdes, Miranda C. N. Cheng
arXiv:2606. 09950v1 Announce Type: new Abstract: Averaging a neural network over its random parameters and marginalizing a Gaussian sector are the same operation, the Schur complement of the eliminated block, and when that block is closed it returns a covariance and its inverse.
By Jin Lei
The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that generalization error diverges when the number of model parameters approaches a critical value from below.
arXiv:2608. 19331v1 Announce Type: cross Abstract: We modify the NN/QFT duality [1] to incorporate the layerwise permutation symmetry of the network, resulting in a $(0\!
By Ro Jefferson, Shradha Ramakrishnan
arXiv:2609.39768v1 Announce Type: cross
Abstract: Certifying a deployed neural network raises decision problems that the verification literature has not classified: whether the model carries a backdo...
By Adrian Wurm
arXiv:2607. 13749v1 Announce Type: new Abstract: Neural networks trained on modular arithmetic exhibit grokking, a delayed transition from memorisation to generalisation known to depend on model capacity: too little and the network memorises slowly or not at all, too much and it generalises almost immediately.
By Chon-Fai Kam, Xavier Cadet, Miloud Bessafi, Frederic Cadet
arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
arXiv:2609.15355v1 Announce Type: cross
Abstract: We study the uniform approximation of smooth scalar-valued functionals on an infinite-dimensional separable Hilbert space by deep ReLU neural network...
By Shuhao Jiao
arXiv:2607. 18930v1 Announce Type: cross Abstract: The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous univariate function on a compact domain to arbitrary error.
By Anuragine S A, Prem Jagadeesan
arXiv:2610.00420v1 Announce Type: new
Abstract: A weight space network (or metanetwork) takes the weights of another neural network as input and predicts properties of it. Most prior work trains such...
By Yuxin Ma, Adir Dayan, Yam Eitan, Haggai Maron, Soledad Villar