The paper introduces TPR-Attention, an attention mechanism that operates over tensor‑product representations to embed structured inductive bias into deep learning models. Experiments on compositional tasks demonstrate that TPR‑Attention outperforms existing architectural components in achieving combinatorial generalization. The results suggest that incorporating explicit compositional structure into neural attention can improve systematic generalization.
By Melisa Civeleko\u{g}lu, Isabeau Pr\'emont-Schwarz
arXiv:2606. 10913v1 Announce Type: new Abstract: We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks.
By Jakob Galley, Vahid Shahverdi, Axel Flinth
arXiv:2606. 00130v1 Announce Type: cross Abstract: We study Automatically Differentiable Nonlinear Tensor Networks (ADNTNs), a family of structured weight generators whose compact core tensors are trained end-to-end by reverse-mode automatic differentiation (AD).
By Andrzej Cichocki, Michal Wietczak
Systematic generalization remains a significant challenge in deep learning. In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortle...
arXiv:2606. 25975v1 Announce Type: new Abstract: Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models.
By Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko, Sergei Kudriashov, Maxim Rakhuba
arXiv:2606. 06772v1 Announce Type: cross Abstract: Understanding the generalization performance of over-parameterized neural networks has become a central topic in deep learning theory.
By Junyu Zhou, Puyu Wang, Yunwen Lei, Marius Kloft, Yiming Ying