arXiv AI

Mixtures of Neural Operators Reduce Active Complexity in Operator Learning

arXiv:2404. 09101v3 Announce Type: replace-cross Abstract: Operator-learning systems are not governed solely by total parameter count; for one query, the relevant bottleneck can be the model that must be loaded and evaluated.

arXiv Machine Learning
Sep 4

Towards a Statistical Understanding of Mixture-of-Experts

The paper presents a statistical framework for Mixture-of-Experts (MoE) models, treating them as localized aggregation systems. It derives oracle risk bounds that separate approximation, expert‑learning, and router‑estimation errors for both dense and sparse routing with evolving experts. The authors also analyze how sparse Top‑K routing balances computational cost with performance, interpret gating geometrically, and explain how shared experts can capture common predictive structure while allowing routed experts to focus on local residuals.

By Siyuan He, Bokai Yang, Jie Hu, Ziwen Gao, Yuhong Yang
arXiv Machine Learning
Aug 18

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

arXiv:2608. 15982v1 Announce Type: new Abstract: We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces.

By Mahdi Mohammadigohari, Thomas Borsani, Giuseppe Di Fatta
Hugging Face Trending Papers
Aug 17

Operator-Theoretic Generalization Bounds for Multitask Deep Learning

We develop operator-theoretic generalization bounds for deep multi-output function classes by representing network layers as Koopman composition operators on vector-valued reproducing kernel Hilbert spaces. In vector-valued Sobolev RKHSs, we derive Rademacher complexity bounds for invertible and width-expanding injective architectures.

arXiv Machine Learning
Jul 16

New universal operator approximation theorem for encoder-decoder architectures

arXiv:2503. 24092v2 Announce Type: replace-cross Abstract: Motivated by the rapidly growing field of mathematics for operator approximation with neural networks, we present a novel universal operator approximation theorem for broad classes of encoder-decoder architectures and a wide range of input and output spaces.

By Janek G\"odeke, Pascal Fernsel
arXiv AI
Sep 18

L2R: Low-Rank and Lipschitz-Controlled Routing for Mixture-of-Experts

The paper introduces L2R, a routing framework for Mixture-of-Experts models that reshapes the routing space into a shared low‑rank latent space and employs Saturated Inner‑Product Scoring to control Lipschitz behavior, resulting in smoother and more stable routing geometry. It also adds a parameter‑efficient multi‑anchor routing mechanism to increase expert expressiveness. Experiments on an OLMoE‑based language model and a ViT‑based ImageNet setting demonstrate improved overall performance and better routing geometry and expert discrimination.

By Minghao Yang, Ren Togo, Guang Li, Takahiro Ogawa, Miki Haseyama
Hugging Face Trending Papers
Jul 7

On Explicit Super-Expressive Approximation for Neural Networks

In this work, we investigate the fixed-architecture neural network approximation with explicit parameter bounds and elementary activations. While prior work demonstrated super-expressive approximation using fixed-size networks, they lack quantitative and non-asymptotic characterizations of parameter magnitude with respect to the approximation error.