arXiv Machine Learning By Travis Pence, Daisuke Yamada, Vikas Singh

Recursive Binding on a Budget: Subspace Carving in Order-p Tensor Memories

Read the original on arXiv Machine Learning →

arXiv:2606. 11391v1 Announce Type: new Abstract: Tensor Product Representations provide the structural fidelity required for symbolic reasoning in models but suffer from exponential dimensionality growth when encoding deep recursive structures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 11

Composing Linear Layers from Irreducibles

arXiv:2507. 11688v4 Announce Type: replace Abstract: Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood.

By Travis Pence, Daisuke Yamada, Vikas Singh
arXiv Machine Learning
Sep 14

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

RunningTensor generalizes linear attention and state‑space models by extending the recurrent memory from a second‑order tensor (matrix) to an order‑o tensor. The memory is updated via a rank‑1 outer product and read by contracting with o‑1 vector queries, with order‑2 recovering linear attention. Experiments on synthetic associative recall and real language tasks show that RunningTensor improves memory capacity from O(W²) to O(Wᵒ) and outperforms existing baselines.

By Luca Herranz-Celotti, Vincent Guigue
arXiv AI
6d ago

DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep Learning

DanLing NestedTensor is a PyTorch tensor abstraction that embeds multi‑ragged structure directly into the tensor, allowing packed values to carry partition information and logical dimension order. This design enables broadcasting, feature transformations, and reductions to automatically respect ragged axes while preserving the same representation through autograd and both eager and compiled execution. Benchmarks on an A100 show significant speedups—up to 3.39× over padding for BERT models and 2.40–4.32× for a Pairformer‑style workload—while dramatically reducing peak memory usage.

By Zhiyuan Chen