arXiv Machine Learning By Niccol\`o Ciolli, Anders Vestergaard N{\o}rskov, Michael Kastoryano, Petr Taborsky, Morten M{\o}rup

(MPO)$^2$: Multivariate Polynomial Optimization based on Matrix Product Operators

Read the original on arXiv Machine Learning →

arXiv:2607. 15916v1 Announce Type: new Abstract: Central to machine learning and signal processing is the ability to perform universal function approximation and learn complex input-output relationships from limited numbers of observations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices

The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.

By Ali Aliev, Maxim Rakhuba
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv Machine Learning
Aug 26

Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation

Polynomial-Augmented Neural Networks (PANNs) merge deep neural networks with polynomial expansions to leverage the flexibility of DNNs and the rapid convergence of polynomials. The architecture introduces orthogonality constraints, basis pruning, and polynomial preconditioning to stabilize training and improve accuracy across diverse problems. Experiments show that PANNs outperform both pure DNNs and polynomial methods in approximating smooth and limited‑smoothness functions, as well as in solving partial differential equations.

By Madison Cooley, Shandian Zhe, Robert M. Kirby, Varun Shankar