arXiv Machine Learning

The Spectral Neuron

arXiv:2608. 08003v1 Announce Type: cross Abstract: As machine learned models increase in complexity and expressive power, features of simpler models, such as interpretability and control over the shape of the modeled function are lost.

arXiv AI
4d ago

Neural networks for spectral optimization

arXiv:2609.36047v1 Announce Type: cross Abstract: Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional....

By Alexis de Villeroch\'e, Beniamin Bogosel, St\'ephane Breuils, Dorin Bucur, Jacques-Olivier Lachaud
arXiv Machine Learning
Sep 3

Learning and extrapolating scale-invariant processes

The paper investigates how machine learning models can regress scale‑free processes, such as earthquakes or avalanches, focusing on predicting rare, large events that require extrapolation. It studies two self‑similar systems: a 2‑dimensional fractional Gaussian field and the Abelian sandpile model. Experiments compare existing architectures (U‑net, Riesz network) with new proposals (wavelet‑based Graph Neural Network, Fourier embedding, Fourier‑Mellin Neural Operator) to identify spectral bias and coarse‑graining challenges and suggest inductive biases to address them.

By Anaclara Alvez-Canepa, Cyril Furtlehner, Fran\c{c}ois P. Landes
arXiv Machine Learning
Sep 10

Observability conditions for neural state-space models with eigenvalues and their roots of unity

The paper investigates observability in neural state‑space models, particularly the Mamba architecture, using tools from ordinary differential equations and control theory. It introduces several strategies—based on eigenvalues, roots of unity, permutations, Fourier transforms, and Vandermonde matrices—to enforce observability in high‑dimensional, learnable hidden states while maintaining computational efficiency. The authors also present a shared‑parameter construction for Mamba and a training algorithm that satisfies a Robbins‑Monro condition, contrasting it with classical procedures that fail to meet contraction requirements.

By Andrew Gracyk
arXiv AI
Sep 15

L-Lipschitz Gershgorin ResNet Network

The paper introduces a method for constructing L-Lipschitz deep residual networks (ResNets) using a Linear Matrix Inequality (LMI) framework. By reformulating the ResNet architecture as a pseudo-tridiagonal LMI and applying the Gershgorin circle theorem, the authors derive closed‑form constraints on network parameters that guarantee Lipschitz continuity. The work also presents a compositional framework for handling recursive systems in hierarchical architectures, while noting that the Gershgorin-based approximations can over‑constrain the system, reducing expressive capacity.

By Marius F. R. Juston, William R. Norris, Dustin Nottage, Ahmet Soylemezoglu
arXiv Machine Learning
Aug 28

Cone Extended Rayleigh Quotients for Directed Graph Learning: Minimax Spectral Certificates, Sensitivity, and Adaptive Control

The paper introduces a learning-oriented framework for spectral certification, sensitivity analysis, and adaptive control of directed graph learning models using cone extended Rayleigh quotients. It provides computable cone bounds and differentiable soft-min/max surrogates that enable rigorous one-sided spectral bounds without requiring symmetry or cone preservation. Experiments on directed networks, including the Cora citation graph, demonstrate that adaptive sensitivity recomputation can significantly reduce spectral levels while preserving test accuracy.

By Yavdat Sh. Il'yasov, Nur F. Valeev
arXiv Machine Learning
Jun 17

Eigen-Spike Emergence and Quadratic Equivalents for Conjugate Kernels on Nonlinearly Separable Data

arXiv:2605. 29669v2 Announce Type: replace-cross Abstract: Recent work in random matrix theory (RMT) has developed the notion of deterministic equivalents: typically linear surrogate models that approximate the spectral behavior of large nonlinear random matrices, such as nonlinear feature maps in neural networks (NNs).

By Collin Cranston, Zhichao Wang, Todd Kemp, Michael W. Mahoney
arXiv Machine Learning
1d ago

Neural scaling laws and evolution of learnable activation functions of Kolmogorov-Arnold networks

The paper studies Kolmogorov‑Arnold Networks (KANs), a neural architecture that treats activation functions as learnable components, offering improved interpretability for scientific applications. It investigates how KANs scale with dataset size on image classification tasks (MNIST, Fashion‑MNIST) and a magnetic‑parameter regression task, revealing a broken neural scaling law that transitions from a faster to a slower decay of test loss as data grows. The authors also analyze how the learned activation functions evolve from simple linear approximations to more complex, interpretable symbolic forms as more data is provided.

By Tilen Cadez, Sanghoon Lee, Kyoung-Min Kim