Vectorizer is a source‑to‑source tool that transforms NumPy programs containing explicit loops into vectorized array operations by rewriting loop bodies from the inside out. It uses array shapes and dataflow analysis to guide a set of rewrite rules that are correct by construction, achieving fast transformations—averaging 0.53 seconds per benchmark. In tests on 150 benchmarks, Vectorizer successfully vectorized 142 directly and 2 with minor edits, producing code that runs on average 74.83× faster than the original loop‑based implementations.
arXiv:2602. 02759v3 Announce Type: replace-cross Abstract: Despite the ubiquity of multiway data across scientific domains, there are few performant and user-friendly methods that fit non-standard nonnegative tensor factorization models tailored to the data at-hand.
By John Hood, Aaron Schein
arXiv:2409. 17502v2 Announce Type: replace Abstract: Broadcast operations are widely used in scientific computing libraries, yet their mathematical formulation is often implicit and inconsistently represented in machine learning literature.
By Yusuke Matsui, Tatsuya Yokota
arXiv:2608.24738v1 Announce Type: new
Abstract: Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e....
By Kai Zhao
Morphological transforms are long-standing tools for shape and mask processing, but the de facto reference implementation in the Python ecosystem, i.e. scipy.ndimage, is CPU-only, single-array, and th...
arXiv:2608. 17135v1 Announce Type: cross Abstract: Tensor networks are powerful formats for compressing large-scale data.
By Xiao Wang, Tomohiro Hashizume, Pia Siegl, Dieter Jaksch
arXiv:2604. 09558v2 Announce Type: replace-cross Abstract: With the widening gap between compute and memory operation latencies, data movement optimizations have become increasingly important for DNN compilation.
By Muyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta, Taeksang Kim, Xin Xu, Janardhan Kulkarni, Ofer Dekel, Vikram Adve, Charith Mendis
DanLing NestedTensor is a PyTorch tensor abstraction that embeds multi‑ragged structure directly into the tensor, allowing packed values to carry partition information and logical dimension order. This design enables broadcasting, feature transformations, and reductions to automatically respect ragged axes while preserving the same representation through autograd and both eager and compiled execution. Benchmarks on an A100 show significant speedups—up to 3.39× over padding for BERT models and 2.40–4.32× for a Pairformer‑style workload—while dramatically reducing peak memory usage.
By Zhiyuan Chen
arXiv:2601. 16622v2 Announce Type: replace-cross Abstract: Equivariant Graph Neural Networks (EGNNs) have become a widely used approach for modeling 3D atomistic systems.
By Lin Huang, Chengxiang Huang, Ziang Wang, Yiyue Du, Chu Wang, Haocheng Lu, Yunyang Li, Xiaoli Liu, Arthur Jiang, Jia Zhang
arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.
By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.
By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki