arXiv Machine Learning

On the robustness of noisy solutions in non-convex neural networks

arXiv:2607. 27000v1 Announce Type: cross Abstract: Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare.

arXiv Machine Learning
Aug 27

Thermodynamic cost of inference and learning in physical neural networks

The paper investigates the thermodynamic cost of inference and learning in physical neural networks. It shows that quasi‑static inference requires no work, while finite‑speed inference incurs work bounded by the Wasserstein‑2 distance between thermal states, roughly $k_B T$ per dimension of the widest layer. Learning, however, has an irreducible cost of a few $k_B T$ per parameter, independent of speed, indicating that memory dominates the thermodynamic price.

By Alexei V. Tkachenko
Hugging Face Trending Papers
Aug 3

Convex Neural Energy Elements: Monolithic Finite-Element Assembly of Geometry-Parameterized Neural Operators with Stability and Error Guarantees

Extending the neural-operator element method from individually trained, fixed-geometry neural elements to a library of reusable, geometry-parameterized element types fails structurally: a field-predicting operator trained by value regression induces an energy whose assembled Hessian is indefinite, and Newton converges to spurious minima (247% error) even with 1%-accurate field predictions. We introduce convex neural energy elements: each element exports a scalar energy E(g,U), architecturally convex in its boundary degrees of freedom U and smoothly parameterized by its geometry g, realized as a hypernetwork-generated positive-semidefinite quadratic form (an input-convex correction is reserved for non-quadratic physics).

arXiv Machine Learning
Aug 4

Convex Neural Energy Elements: Monolithic Finite-Element Assembly of Geometry-Parameterized Neural Operators with Stability and Error Guarantees

arXiv:2608. 02036v1 Announce Type: new Abstract: Extending the neural-operator element method from individually trained, fixed-geometry neural elements to a library of reusable, geometry-parameterized element types fails structurally: a field-predicting operator trained by value regression induces an energy whose assembled Hessian is indefinite, and Newton converges to spurious minima (247% error) even with 1%-accurate field predictions.

By Hongyue Jiang, Jianjiang Zhan, Chenzhuo Zhang, Fan Wang
arXiv AI
Sep 24

Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon

The paper presents a near-complete, nonasymptotic generalization theory for multilayer neural networks using path regularization, applicable to broad Lipschitz loss functions without requiring bounded loss or extreme network hyperparameters. It provides an explicit upper bound that addresses approximation rates in generalized Barron spaces and demonstrates the double descent phenomenon for ReLU networks. The authors claim near-minimax optimality for regression problems and plan to establish matching lower bounds in future work.

By Hao Yu