The paper introduces a quantum tensor network learning framework that employs matrix product states (MPS) as a machine‑learning architecture, adding a global normalization condition to interpret the MPS as a quantum state. It compares two optimization strategies—gradient descent and a DMRG‑adapted method—to identify locally optimal tensors and evaluates their effectiveness.
By Gustav J L J\"ager, Martin B Plenio, Hans-Martin Rieser
arXiv:2502. 09928v2 Announce Type: replace-cross Abstract: Originating in quantum physics, tensor networks (TNs) have been widely adopted as exponential machines and parametric decomposers for recognition tasks.
By Chang Nie
arXiv:2608. 17135v1 Announce Type: cross Abstract: Tensor networks are powerful formats for compressing large-scale data.
By Xiao Wang, Tomohiro Hashizume, Pia Siegl, Dieter Jaksch
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.
arXiv:2606. 08565v1 Announce Type: cross Abstract: Tensor networks provide efficient representations for compressing large neural networks.
By Toshiaki Koike-Akino, Jing Liu, Ye Wang
arXiv:2608.21700v1 Announce Type: cross
Abstract: Continuous-time flow and diffusion models are widely used across many application domains, from large-scale deployment in computer vision and protein...
By Nathan X. Kodama, L. Andrew Wray, Sam Cochran, Chad Rigetti, Shravan Veerapaneni, Michael J. Keiser
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
By Yushun Zhang
arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.
By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv:2606. 00130v2 Announce Type: replace-cross Abstract: Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters.
By Andrzej Cichocki, Michal Wietczak
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics.
arXiv:2606. 00130v1 Announce Type: cross Abstract: We study Automatically Differentiable Nonlinear Tensor Networks (ADNTNs), a family of structured weight generators whose compact core tensors are trained end-to-end by reverse-mode automatic differentiation (AD).
By Andrzej Cichocki, Michal Wietczak
arXiv:2604. 15645v2 Announce Type: replace Abstract: We present QPINNACLE, an open-source computational framework for physics-informed neural networks (PINNs) that integrates modern training strategies, multi-GPU acceleration, and hybrid quantum-classical architectures within a unified modular workflow.
By Ziv Chen, Hemanth Chandravamsi, Shimon Pisnoy, Aaron Goldgewert, Gal Shaviner, Boris Shragner, Steven H. Frankel