The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.
By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us
The paper presents a formal analysis of the quotient geometry of tree tensor networks (TTNs) and introduces efficient first- and second-order optimization algorithms that leverage this geometry. It also develops a backpropagation method for training TTNs in a kernel learning context. Numerical experiments on a digit classification task demonstrate a tradeoff between two horizontal distributions: one provides clearer geometric insights, while the other yields more efficient algorithms.
By Marius Willner, Marco Trenti, Dirk Lebiedz
arXiv:2410. 17397v2 Announce Type: replace-cross Abstract: We introduce a framework for seamlessly integrating quantum computing into pretrained large language models (LLMs).
By Borja Aizpurua, Fernando Loren, Saeed S. Jahromi, Sukhbinder Singh, Roman Orus
arXiv:2606. 00130v2 Announce Type: replace-cross Abstract: Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters.
By Andrzej Cichocki, Michal Wietczak
arXiv:2606. 00130v1 Announce Type: cross Abstract: We study Automatically Differentiable Nonlinear Tensor Networks (ADNTNs), a family of structured weight generators whose compact core tensors are trained end-to-end by reverse-mode automatic differentiation (AD).
By Andrzej Cichocki, Michal Wietczak
arXiv:2609.00870v1 Announce Type: cross
Abstract: Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning. We derive stochastic Riemannian optim...
By Marius Willner, Maximilian Scharf, Andr\'e Uschmajew, Timo Felser, Marco Trenti
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
arXiv:2607. 18074v1 Announce Type: new Abstract: Equivariant graph neural networks repeatedly apply edge-conditioned tensor-product convolutions over graph edges.
By Vladimir Choro\v{s}ajev, C\'edric B\'eny
The paper introduces a quantum tensor network learning framework that employs matrix product states (MPS) as a machine‑learning architecture, adding a global normalization condition to interpret the MPS as a quantum state. It compares two optimization strategies—gradient descent and a DMRG‑adapted method—to identify locally optimal tensors and evaluates their effectiveness.
By Gustav J L J\"ager, Martin B Plenio, Hans-Martin Rieser
This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.
By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
arXiv:2601. 16622v2 Announce Type: replace-cross Abstract: Equivariant Graph Neural Networks (EGNNs) have become a widely used approach for modeling 3D atomistic systems.
By Lin Huang, Chengxiang Huang, Ziang Wang, Yiyue Du, Chu Wang, Haocheng Lu, Yunyang Li, Xiaoli Liu, Arthur Jiang, Jia Zhang
arXiv:2608. 19789v1 Announce Type: new Abstract: Developed as a workhorse for classical simulations of quantum algorithms and quantum many-body systems, Tensor Network methods have entered the scientific mainstream in quantum physics.
By Michal A. Sterzel, Marko J. Ran\v{c}i\'c