arXiv AI By Michael L. Roberts, Carlos Zapata Carratal\'a. Nicholas J. Cooper, Lijun Chen, Fran\c{c}ois G. Meyer, Danna Gurari

Higher Structures in Deep Learning

Read the original on arXiv AI →

The article offers an expository overview of why higher‑arity tensor operations matter for deep learning. It presents an empirical study of higher‑arity phenomena in trained neural networks, introduces a hypergraphical generalization of the multilayer perceptron, and examines links to evolutionary algorithms. The paper concludes by outlining promising future research directions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

TPR-Attention for Combinatorial Generalization

The paper introduces TPR-Attention, an attention mechanism that operates over tensor‑product representations to embed structured inductive bias into deep learning models. Experiments on compositional tasks demonstrate that TPR‑Attention outperforms existing architectural components in achieving combinatorial generalization. The results suggest that incorporating explicit compositional structure into neural attention can improve systematic generalization.

By Melisa Civeleko\u{g}lu, Isabeau Pr\'emont-Schwarz
arXiv Machine Learning
Jun 25

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.

By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba