arXiv Machine Learning

Tensor Data Scattering and the Impossibility of Slicing Theorem

arXiv:2012. 01982v3 Announce Type: replace Abstract: This paper proposes a standard way to represent sparse tensors.

arXiv Machine Learning
Jul 9

VTC: DNN Compilation with Virtual Tensors for Data Movement Elimination

arXiv:2604. 09558v2 Announce Type: replace-cross Abstract: With the widening gap between compute and memory operation latencies, data movement optimizations have become increasingly important for DNN compilation.

By Muyan Hu, Ahan Gupta, Jiachen Yuan, Vima Gupta, Taeksang Kim, Xin Xu, Janardhan Kulkarni, Ofer Dekel, Vikram Adve, Charith Mendis
arXiv Machine Learning
Aug 31

A Deeper Analysis of Block-Sparse Featurizers

The paper investigates the block-sparse featurizer (BSF), a model that uses small subspaces as atomic units instead of single directions, aiming to capture features on low-dimensional manifolds common in vision. It identifies that BSF still exhibits classic sparse autoencoder failure modes such as feature splitting and composition. The authors propose architectural improvements, notably a Tournament Top‑K selection rule, which markedly reduces feature splitting, and they extend the block concept to a crosscoder framework.

By Alexandru-Iulius Jerpelea, Amith Ananthram
arXiv Machine Learning
Aug 28

ReLATE: Accelerating Tensor Decomposition via Safe and Efficient Learning of Sparse Encodings

ReLATE is a reinforcement‑learned framework that automatically discovers safe and efficient sparse encodings for tensor decomposition, eliminating the need for expert‑designed formats. It combines model‑free and model‑based learning, elastic training, rule‑driven action masking, and dynamics‑informed filtering to guarantee correct encoding with bounded execution time, even early in training. After offline training, ReLATE deploys the optimal encoding with negligible overhead and achieves up to 2× speedups over the best expert‑designed format, with a geometric‑mean speedup of 1.38–1.41× on diverse real‑world sparse tensors.

By Ahmed E. Helal, Fabio Checconi, Jan Laukemann, Yongseok Soh, Jesmin Jahan Tithi, Fabrizio Petrini, Jee Choi
Hugging Face Trending Papers
Aug 3

CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization

Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.

arXiv AI
6d ago

DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning

DanLing NestedTensor is a PyTorch tensor abstraction that embeds multi‑ragged structure directly into the tensor, allowing packed values to carry partition information and logical dimension order. This design enables broadcasting, feature transformations, and reductions to automatically respect ragged axes while preserving the same representation through autograd and both eager and compiled execution. Benchmarks on an A100 show significant speedups—up to 3.39× over padding for BERT models and 2.40–4.32× for a Pairformer‑style workload—while dramatically reducing peak memory usage.

By Zhiyuan Chen