arXiv:2606. 00130v1 Announce Type: cross Abstract: We study Automatically Differentiable Nonlinear Tensor Networks (ADNTNs), a family of structured weight generators whose compact core tensors are trained end-to-end by reverse-mode automatic differentiation (AD).
By Andrzej Cichocki, Michal Wietczak
arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.
By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics.
arXiv:2608. 17135v1 Announce Type: cross Abstract: Tensor networks are powerful formats for compressing large-scale data.
By Xiao Wang, Tomohiro Hashizume, Pia Siegl, Dieter Jaksch
arXiv:2606. 08565v1 Announce Type: cross Abstract: Tensor networks provide efficient representations for compressing large neural networks.
By Toshiaki Koike-Akino, Jing Liu, Ye Wang
arXiv:2609.00870v1 Announce Type: cross
Abstract: Tensor networks, originally developed for quantum many-body physics, are promising models for machine learning. We derive stochastic Riemannian optim...
By Marius Willner, Maximilian Scharf, Andr\'e Uschmajew, Timo Felser, Marco Trenti
arXiv:2606. 03465v1 Announce Type: cross Abstract: Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints.
By Artur Zagitov, Alexander Miasnikov, Maxim Krutikov, Vladimir Aletov, Gleb Molodtsov, Nail Bashirov, Artem Tsedenov, Aleksandr Beznosikov
arXiv:2608. 01633v1 Announce Type: new Abstract: Large language models (LLMs) enable neural architecture search (NAS) directly over executable neural network programs.
By Zhen Liu, Wanqi Zhou, Shuanghao Bai, Yuhan Liu, Jinjun Wang, Jingwen Fu
Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged as a promising direction, offering compact parameterizations well suited to Transformer weight structures.
The article offers an expository overview of why higher‑arity tensor operations matter for deep learning. It presents an empirical study of higher‑arity phenomena in trained neural networks, introduces a hypergraphical generalization of the multilayer perceptron, and examines links to evolutionary algorithms. The paper concludes by outlining promising future research directions.
By Michael L. Roberts, Carlos Zapata Carratal\'a. Nicholas J. Cooper, Lijun Chen, Fran\c{c}ois G. Meyer, Danna Gurari
arXiv:2606. 00130v2 Announce Type: replace-cross Abstract: Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters.
By Andrzej Cichocki, Michal Wietczak
The paper introduces a quantum tensor network learning framework that employs matrix product states (MPS) as a machine‑learning architecture, adding a global normalization condition to interpret the MPS as a quantum state. It compares two optimization strategies—gradient descent and a DMRG‑adapted method—to identify locally optimal tensors and evaluates their effectiveness.
By Gustav J L J\"ager, Martin B Plenio, Hans-Martin Rieser