arXiv Machine Learning

Near-Universal Multiplicative Updates for Nonnegative Einsum Factorization

arXiv:2602. 02759v3 Announce Type: replace-cross Abstract: Despite the ubiquity of multiway data across scientific domains, there are few performant and user-friendly methods that fit non-standard nonnegative tensor factorization models tailored to the data at-hand.

arXiv AI
Sep 1

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.

By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
arXiv Machine Learning
Aug 6

E$^2$M: Double Bounded $\alpha$-Divergence Optimization for Tensor-based Discrete Density Estimation

arXiv:2405. 18220v4 Announce Type: replace-cross Abstract: Tensor-based discrete density estimation requires flexible modeling and proper divergence criteria to enable effective learning; however, traditional approaches using $\alpha$-divergence face analytical challenges due to the $\alpha$-power terms in the objective function, which hinder the derivation of closed-form update rules.

By Kazu Ghalamkari, Jesper L{\o}ve Hinrich, Morten M{\o}rup
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
Hugging Face Trending Papers
Sep 2

Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing

The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both run‑time and root‑sum‑of‑errors accuracy for predicting drug effects.

arXiv Machine Learning
Sep 4

Coupled Tensor-Tensor Completion Method with Applications in Drug Repurposing

The paper introduces Coupled Tensor‑Tensor Completion (CTTC), a new framework that incorporates side information in tensor form to enhance tensor completion tasks. CTTC leverages hidden connections among multimodal tensors and is grounded in distance metric learning and group theory. Experiments on the DTD and LINCS datasets show that CTTC outperforms existing methods such as HaLRTC, CTRC, Cell, and NTDDR in both runtime and root‑sum‑of‑squares error for drug effect prediction.

By Maryam Bagherian, Albert Hung, Ivo Dinov, Joshua Welch