arXiv:2206. 04359v3 Announce Type: replace Abstract: One of the fundamental challenges in the deep learning community is to theoretically understand how well a deep neural network generalizes to unseen data.
By Chengli Tan, Jiangshe Zhang, Junmin Liu, Yihong Gong
arXiv:2606. 18071v1 Announce Type: cross Abstract: Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless noising.
By Yusen Jia, Bingyan Han
arXiv:2607. 15505v1 Announce Type: cross Abstract: Fractional gradient descent (FGD) incorporates long-range memory through Caputo-type operators and has been shown to improve stability in ill-conditioned and nonconvex optimization problems.
By Hwanseo Lee, Junseo Lee, Hyunju Kim
arXiv:2608. 14636v1 Announce Type: cross Abstract: Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training.
By Sebastian Raubitzek, Georg Goldenits, Sebastian Schrittwieser, Philip K\"onig, Kevin Mallinger
arXiv:2607. 29038v1 Announce Type: new Abstract: Fractional scientific machine learning requires numerical operators that can be differentiated, batched, accelerated, and composed with neural networks.
By Ning Hu, Haitao Duan, Shuqun Li, Chuyang Hu
arXiv:2608. 12879v1 Announce Type: new Abstract: Fractional partial differential equations describe nonlocal dynamics, but discovering them from noisy data is difficult because fractional differentiation amplifies high-frequency measurement noise and the derivative orders are unknown.
By Pongpisit Thanasutives, Yoshinobu Kawahara
The paper introduces the fractional Laplace neural operator (fLNO), a neural operator that embeds Volterra resolvent structures with non‑rational Laplace symbols into learned maps. It demonstrates that a single graph‑spectral layer can exactly represent the full linear Volterra solution for commuting excitation–Laplacian pairs, and establishes limits on the expressivity of finite rational realizations, showing they cannot capture non‑integer critical asymptotics. The authors also provide trainable parametrizations that enforce stability margins, a graphon‑transfer theorem, and empirical results on benchmark data, Chilean aftershock sequences, and renewal models, highlighting the fLNO’s ability to recover branching coordinates with few parameters while maintaining stability.
By Mauricio Herrera-Mar\'in
Spiking Neural Networks (SNNs) are well-regarded for their biological plausibility and energy efficiency in processing sequential data. However, dominant SNN architectures typically rely on first-order Ordinary Differential Equations (ODEs) to govern neuronal state transitions.
arXiv:2609. 18127v1 Announce Type: new Abstract: Many real-world processes exhibit long-range dependence, where the current state depends on a slowly decaying trace of past states rather than on the most recent state alone.
By Xiaole Zhang, Ziyi Zhang, Zehao Zhao, Stephen Tu, Guannan Qu, Yorie Nakahira, Paul Bogdan
arXiv:2609.36314v1 Announce Type: new
Abstract: State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for...
By Ivan Kobyzev, Abbas Ghaddar, Ali Nasiri-Sarvi, Lifeng Shang, Yufei Cui
The paper investigates how deep residual networks behave when their initial weights are correlated across layers. It confirms a conjecture that such correlated initializations interpolate between a Brownian stochastic differential equation (for independent weights) and an ordinary differential equation (for perfectly correlated weights). By applying a feature function to a stationary Gaussian sequence with regularly varying correlation, the authors prove that a unique critical scaling exists, leading the infinite‑depth limit to a Young differential equation driven by a Hermite process, which reduces to fractional Brownian motion when the feature function has Hermite rank one. The study shows that the correlation structure and Hermite rank of the initialization uniquely determine the critical scaling and asymptotic limit, making them meaningful hyperparameters in the asymptotic regime, whereas finite‑variance i.i.d. initialization always yields a Brownian driver regardless of distribution.
By Felix Benning, Ivan Nourdin, Giovanni Peccati
The paper introduces Chernoff-neural operators, a class of neural operators that can universally approximate Chernoff-type one-step operators for strongly continuous convex monotone semigroups. A universal approximation theorem is proved, and stability estimates in weighted Hölder spaces allow the one-step error to propagate, yielding universal approximation of the entire semigroup. The authors also define envelope-neural operators for envelope semigroups, providing quantitative approximation rates, and demonstrate the approach on numerical examples from nonlinear PDEs, stochastic optimal control, and uncertain stochastic processes.
By Jonas Blessing, Philipp Schmocker, Alessandro Sgarabottolo