Beyond Additive Decompositions: Interpretability Through Separability
arXiv:2605. 31200v2 Announce Type: replace Abstract: Interpretable machine learning requires models that are accurate and structurally faithful to the data.
arXiv:2607. 15916v1 Announce Type: new Abstract: Central to machine learning and signal processing is the ability to perform universal function approximation and learn complex input-output relationships from limited numbers of observations.
arXiv:2605. 31200v2 Announce Type: replace Abstract: Interpretable machine learning requires models that are accurate and structurally faithful to the data.
The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.
arXiv:2608. 07043v1 Announce Type: cross Abstract: Developing nonlinear models that are both expressive and computationally efficient remains a challenge in machine learning and nonlinear system identification.
arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.
arXiv:2409. 17502v2 Announce Type: replace Abstract: Broadcast operations are widely used in scientific computing libraries, yet their mathematical formulation is often implicit and inconsistently represented in machine learning literature.
Polynomial-Augmented Neural Networks (PANNs) merge deep neural networks with polynomial expansions to leverage the flexibility of DNNs and the rapid convergence of polynomials. The architecture introduces orthogonality constraints, basis pruning, and polynomial preconditioning to stabilize training and improve accuracy across diverse problems. Experiments show that PANNs outperform both pure DNNs and polynomial methods in approximating smooth and limited‑smoothness functions, as well as in solving partial differential equations.
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics.
arXiv:2608. 10351v1 Announce Type: new Abstract: In this work we present a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN).
arXiv:2606. 31061v1 Announce Type: cross Abstract: Tensor Train (TT) decomposition is a powerful technique for analyzing high-dimensional data.
arXiv:2606.08560v2 Announce Type: replace-cross Abstract: We adopt the canonical polyadic (CP) decomposition to model high-dimensional tensor time series. Our primary goal is to identify and estimate...
The paper introduces an online framework for functional principal component analysis (FPCA) tailored to multidimensional functional data streams. It models functional principal components with tensor product splines, enforcing smoothness and orthonormality via a penalized approach on a Stiefel manifold. The authors present efficient Riemannian stochastic gradient descent and AdaGrad algorithms, along with a dynamic smoothing parameter tuning strategy based on rolling block validation, and provide asymptotic normality results and confidence intervals for the estimators.
arXiv:2407. 00809v4 Announce Type: replace Abstract: This paper introduces the Kernel Neural Operator (KNO), a provably convergent operator-learning architecture that utilizes compositions of deep kernel-based integral operators for function-space approximation of operators (maps from functions to functions).