arXiv AI

Deep Tree Tensor Networks

arXiv:2502. 09928v2 Announce Type: replace-cross Abstract: Originating in quantum physics, tensor networks (TNs) have been widely adopted as exponential machines and parametric decomposers for recognition tasks.

arXiv AI
Sep 15

Tensorization is a powerful but underexplored tool for compression and interpretability of neural networks

The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.

By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us
arXiv Machine Learning
Sep 23

Riemannian Optimization on Tree Tensor Networks with Application in Machine Learning

The paper presents a formal analysis of the quotient geometry of tree tensor networks (TTNs) and introduces efficient first- and second-order optimization algorithms that leverage this geometry. It also develops a backpropagation method for training TTNs in a kernel learning context. Numerical experiments on a digit classification task demonstrate a tradeoff between two horizontal distributions: one provides clearer geometric insights, while the other yields more efficient algorithms.

By Marius Willner, Marco Trenti, Dirk Lebiedz
arXiv Computer Vision
Sep 2

Stochastic Optimization of Tree Tensor Networks

arXiv:2609.00870v1 Announce Type: cross Abstract: Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning. We derive stochastic Riemannian optim...

By Marius Willner, Maximilian Scharf, Andr\'e Uschmajew, Timo Felser, Marco Trenti
arXiv Machine Learning
Aug 20

Quantum Tensor Network Learning with DMRG

The paper introduces a quantum tensor network learning framework that employs matrix product states (MPS) as a machine‑learning architecture, adding a global normalization condition to interpret the MPS as a quantum state. It compares two optimization strategies—gradient descent and a DMRG‑adapted method—to identify locally optimal tensors and evaluates their effectiveness.

By Gustav J L J\"ager, Martin B Plenio, Hans-Martin Rieser
arXiv AI
Sep 1

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.

By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki