Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged as a promising direction, offering compact parameterizations well suited to Transformer weight structures.
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.
By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
arXiv:2411. 09816v5 Announce Type: replace Abstract: Large neural networks achieve state-of-the-art performance on many tasks, yet their sheer size hinders deployment on resource-constrained devices.
By Cem \"Uy\"uk, Mike Lasby, Mohamed Yassin, Utku Evci, Yani Ioannou
arXiv:2605. 30836v2 Announce Type: replace Abstract: Recent SVD based compression methods for large language models like SVD LLM and Basis Sharing can be unified under one optimization problem.
By Snigdha Chandan Khilar
arXiv:2606. 08347v1 Announce Type: cross Abstract: Modern language models represent text using discrete token-level embeddings, which forces recurring multi-token patterns to be learned implicitly across Transformer layers.
By Wuyang Zhou, Yuxuan Gu, Giorgos Iacovides, Yuning Qiu, Qibin Zhao, Danilo Mandic
arXiv:2506. 09105v3 Announce Type: replace-cross Abstract: We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers.
By Javier Lopez-Piqueres, Pranav Deshpande, Archan Ray, Mattia J. Villani, Marco Pistoia, Niraj Kumar
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.
arXiv:2606. 08565v1 Announce Type: cross Abstract: Tensor networks provide efficient representations for compressing large neural networks.
By Toshiaki Koike-Akino, Jing Liu, Ye Wang
DanLing NestedTensor is a PyTorch tensor abstraction that embeds multi‑ragged structure directly into the tensor, allowing packed values to carry partition information and logical dimension order. This design enables broadcasting, feature transformations, and reductions to automatically respect ragged axes while preserving the same representation through autograd and both eager and compiled execution. Benchmarks on an A100 show significant speedups—up to 3.39× over padding for BERT models and 2.40–4.32× for a Pairformer‑style workload—while dramatically reducing peak memory usage.
By Zhiyuan Chen
arXiv:2607. 25504v1 Announce Type: cross Abstract: Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models.
By Bowen Wang, Chi Zhang, Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini
The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.
By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us