Neural networks for spectral optimization
arXiv:2609.36047v1 Announce Type: cross Abstract: Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional....
arXiv:2608. 08003v1 Announce Type: cross Abstract: As machine learned models increase in complexity and expressive power, features of simpler models, such as interpretability and control over the shape of the modeled function are lost.
arXiv:2609.36047v1 Announce Type: cross Abstract: Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional....
The paper investigates how machine learning models can regress scale‑free processes, such as earthquakes or avalanches, focusing on predicting rare, large events that require extrapolation. It studies two self‑similar systems: a 2‑dimensional fractional Gaussian field and the Abelian sandpile model. Experiments compare existing architectures (U‑net, Riesz network) with new proposals (wavelet‑based Graph Neural Network, Fourier embedding, Fourier‑Mellin Neural Operator) to identify spectral bias and coarse‑graining challenges and suggest inductive biases to address them.
The paper investigates observability in neural state‑space models, particularly the Mamba architecture, using tools from ordinary differential equations and control theory. It introduces several strategies—based on eigenvalues, roots of unity, permutations, Fourier transforms, and Vandermonde matrices—to enforce observability in high‑dimensional, learnable hidden states while maintaining computational efficiency. The authors also present a shared‑parameter construction for Mamba and a training algorithm that satisfies a Robbins‑Monro condition, contrasting it with classical procedures that fail to meet contraction requirements.
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
The paper introduces a method for constructing L-Lipschitz deep residual networks (ResNets) using a Linear Matrix Inequality (LMI) framework. By reformulating the ResNet architecture as a pseudo-tridiagonal LMI and applying the Gershgorin circle theorem, the authors derive closed‑form constraints on network parameters that guarantee Lipschitz continuity. The work also presents a compositional framework for handling recursive systems in hierarchical architectures, while noting that the Gershgorin-based approximations can over‑constrain the system, reducing expressive capacity.
arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
The paper introduces a learning-oriented framework for spectral certification, sensitivity analysis, and adaptive control of directed graph learning models using cone extended Rayleigh quotients. It provides computable cone bounds and differentiable soft-min/max surrogates that enable rigorous one-sided spectral bounds without requiring symmetry or cone preservation. Experiments on directed networks, including the Cora citation graph, demonstrate that adaptive sensitivity recomputation can significantly reduce spectral levels while preserving test accuracy.
arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
arXiv:2609. 23529v1 Announce Type: new Abstract: Neural operators have emerged as powerful surrogates for solving partial differential equations (PDEs), yet their reliability under distribution shift remains a critical barrier to deployment.
arXiv:2605. 29669v2 Announce Type: replace-cross Abstract: Recent work in random matrix theory (RMT) has developed the notion of deterministic equivalents: typically linear surrogate models that approximate the spectral behavior of large nonlinear random matrices, such as nonlinear feature maps in neural networks (NNs).
arXiv:2607. 03347v1 Announce Type: new Abstract: We consider the Multiscale Single-Index Model (MSIM), first introduced in \cite{oymak2021learning}, as a stylized model for hierarchical learning with \emph{scale separation}.
The paper studies Kolmogorov‑Arnold Networks (KANs), a neural architecture that treats activation functions as learnable components, offering improved interpretability for scientific applications. It investigates how KANs scale with dataset size on image classification tasks (MNIST, Fashion‑MNIST) and a magnetic‑parameter regression task, revealing a broken neural scaling law that transitions from a faster to a slower decay of test loss as data grows. The authors also analyze how the learned activation functions evolve from simple linear approximations to more complex, interpretable symbolic forms as more data is provided.