arXiv Machine Learning

Online TT-ALS for Streaming Tensor Decomposition with Incremental Orthogonalization

arXiv:2606. 31061v1 Announce Type: cross Abstract: Tensor Train (TT) decomposition is a powerful technique for analyzing high-dimensional data.

arXiv Machine Learning
Sep 11

Semi-Tensor Product-Based Multi-Term Randomized T-SVD and Its Visual Applications

The paper introduces a new semi‑tensor product for third‑order tensors that relaxes the dimensional constraints of the standard t‑product while preserving the closed‑form nature of T‑SVD. It builds a multi‑term semi‑tensor product singular value decomposition (MSTP‑SVD) that improves low‑rank approximation accuracy, and further accelerates it with randomized projection and power iteration to create the MRSTP‑SVD algorithm. Experiments on image and video compression and completion show that this method balances reconstruction accuracy and computational efficiency.

By Xingchen Xiao (School of Mathematics and Statistics, Southwest University, Chongqing, China), Feng Zhang (School of Mathematics and Statistics, Southwest University, Chongqing, China), Wenjin Qin (School of Mathematics and Statistics, Southwest University, Chongqing, China), Jianjun Wang (School of Mathematics and Statistics, Southwest University, Chongqing, China)
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
Hugging Face Trending Papers
Aug 3

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.

arXiv AI
Sep 10

Mind the Approximation: Fisher-Weighted SVD Compression for ViTs

The paper introduces FACTS, a structured Fisher Approximation for compressing Vision Transformers (ViTs) using Fisher-weighted SVD, which enforces token‑local aggregation while preserving within‑token activation‑gradient dependence. It also presents Constrained Rank Search (CoRS) to optimize layer‑wise rank allocation under a fixed FLOP budget. Experiments on ViTs and hybrid architectures show that FACTS improves accuracy‑efficiency trade‑offs, outperforming the strongest SVD baseline by up to +5.8 percentage points on Swin‑B without requiring finetuning.

By Moritz Thoma, Maximilian Groezinger, Maximilian Forstenh\"ausler, Emad Aghajanzadeh, Ryan Pegoud, Manoj Rohit Vemparala, Pierpaolo Mori, Alexander Frickenstein, Daniel Mueller-Gritschneder, Ulf Schlichtmann