arXiv:2609.36047v1 Announce Type: cross
Abstract: Given a functional dependent on the spectrum of a differential operator, we address the problem of finding a domain which optimizes this functional....
By Alexis de Villeroch\'e, Beniamin Bogosel, St\'ephane Breuils, Dorin Bucur, Jacques-Olivier Lachaud
The paper investigates how machine learning models can regress scale‑free processes, such as earthquakes or avalanches, focusing on predicting rare, large events that require extrapolation. It studies two self‑similar systems: a 2‑dimensional fractional Gaussian field and the Abelian sandpile model. Experiments compare existing architectures (U‑net, Riesz network) with new proposals (wavelet‑based Graph Neural Network, Fourier embedding, Fourier‑Mellin Neural Operator) to identify spectral bias and coarse‑graining challenges and suggest inductive biases to address them.
By Anaclara Alvez-Canepa, Cyril Furtlehner, Fran\c{c}ois P. Landes
The paper investigates observability in neural state‑space models, particularly the Mamba architecture, using tools from ordinary differential equations and control theory. It introduces several strategies—based on eigenvalues, roots of unity, permutations, Fourier transforms, and Vandermonde matrices—to enforce observability in high‑dimensional, learnable hidden states while maintaining computational efficiency. The authors also present a shared‑parameter construction for Mamba and a training algorithm that satisfies a Robbins‑Monro condition, contrasting it with classical procedures that fail to meet contraction requirements.
By Andrew Gracyk
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
The paper introduces a method for constructing L-Lipschitz deep residual networks (ResNets) using a Linear Matrix Inequality (LMI) framework. By reformulating the ResNet architecture as a pseudo-tridiagonal LMI and applying the Gershgorin circle theorem, the authors derive closed‑form constraints on network parameters that guarantee Lipschitz continuity. The work also presents a compositional framework for handling recursive systems in hierarchical architectures, while noting that the Gershgorin-based approximations can over‑constrain the system, reducing expressive capacity.
By Marius F. R. Juston, William R. Norris, Dustin Nottage, Ahmet Soylemezoglu
arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
By Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang