EinSort: Sorting is All We Need for Tensorizing LLM
arXiv:2606. 08565v1 Announce Type: cross Abstract: Tensor networks provide efficient representations for compressing large neural networks.
arXiv:2012. 01982v3 Announce Type: replace Abstract: This paper proposes a standard way to represent sparse tensors.
arXiv:2606. 08565v1 Announce Type: cross Abstract: Tensor networks provide efficient representations for compressing large neural networks.
arXiv:2604. 09558v2 Announce Type: replace-cross Abstract: With the widening gap between compute and memory operation latencies, data movement optimizations have become increasingly important for DNN compilation.
arXiv:2608. 17135v1 Announce Type: cross Abstract: Tensor networks are powerful formats for compressing large-scale data.
arXiv:2606. 03465v1 Announce Type: cross Abstract: Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints.
Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged as a promising direction, offering compact parameterizations well suited to Transformer weight structures.
The paper investigates the block-sparse featurizer (BSF), a model that uses small subspaces as atomic units instead of single directions, aiming to capture features on low-dimensional manifolds common in vision. It identifies that BSF still exhibits classic sparse autoencoder failure modes such as feature splitting and composition. The authors propose architectural improvements, notably a Tournament Top‑K selection rule, which markedly reduces feature splitting, and they extend the block concept to a crosscoder framework.
ReLATE is a reinforcement‑learned framework that automatically discovers safe and efficient sparse encodings for tensor decomposition, eliminating the need for expert‑designed formats. It combines model‑free and model‑based learning, elastic training, rule‑driven action masking, and dynamics‑informed filtering to guarantee correct encoding with bounded execution time, even early in training. After offline training, ReLATE deploys the optimal encoding with negligible overhead and achieves up to 2× speedups over the best expert‑designed format, with a geometric‑mean speedup of 1.38–1.41× on diverse real‑world sparse tensors.
arXiv:2606. 31061v1 Announce Type: cross Abstract: Tensor Train (TT) decomposition is a powerful technique for analyzing high-dimensional data.
arXiv:2607. 06048v1 Announce Type: cross Abstract: We aim to identify scattering network architectures that maximize the separation capacity on data with low intrinsic dimension.
Diffusion Transformers (DiTs) have achieved state-of-the-art (SOTA) performance in visual generative modeling, yet their training remains computationally prohibitive. While the recently proposed Momentum Orthogonalization (Muon) optimizer offers a promising alternative to AdamW, its direct application to DiTs yields suboptimal late-stage convergence.
arXiv:2608. 10351v1 Announce Type: new Abstract: In this work we present a method to accelerate the optimization of learning high dimensional functions using deep neural network (DNN).
DanLing NestedTensor is a PyTorch tensor abstraction that embeds multi‑ragged structure directly into the tensor, allowing packed values to carry partition information and logical dimension order. This design enables broadcasting, feature transformations, and reductions to automatically respect ragged axes while preserving the same representation through autograd and both eager and compiled execution. Benchmarks on an A100 show significant speedups—up to 3.39× over padding for BERT models and 2.40–4.32× for a Pairformer‑style workload—while dramatically reducing peak memory usage.