The paper introduces TPR-Attention, an attention mechanism that operates over tensor‑product representations to embed structured inductive bias into deep learning models. Experiments on compositional tasks demonstrate that TPR‑Attention outperforms existing architectural components in achieving combinatorial generalization. The results suggest that incorporating explicit compositional structure into neural attention can improve systematic generalization.
By Melisa Civeleko\u{g}lu, Isabeau Pr\'emont-Schwarz
arXiv:2601. 04509v2 Announce Type: replace Abstract: Mixed-integer linear programming (MILP) is a foundational framework for combinatorial optimization across science and engineering, but remains hard to solve at scale due to NP-hardness.
By Peixin Huang, Yaoxin Wu, Yining Ma, Cathy Wu, Wei Zhang, Wen Song
arXiv:2603. 00742v2 Announce Type: replace Abstract: While Adam has long been the ubiquitous default optimizer for deep neural networks, Muon has recently seen rapid adoption due to its superior training speed.
By Sara Dragutinovi\'c, Yedi Zhang, Rajesh Ranganath
arXiv:2607. 23634v1 Announce Type: cross Abstract: Attention enables context modeling via query-key scoring with softmax normalization.
By Rui Wang
arXiv:2602. 24264v2 Announce Type: replace-cross Abstract: Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems.
By Arnas Uselis, Andrea Dittadi, Seong Joon Oh
Over the past decade, deep neural networks (DNNs) have achieved remarkable success on complex machine-learning tasks, yet the theoretical foundations of their performance remain incomplete. From a statistical viewpoint, a natural question is: can DNNs attain feature-learning and prediction consistency comparable to that of classical models?
arXiv:2109. 13445v3 Announce Type: replace-cross Abstract: The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not well understood.
By Avi Cooper, Xavier Boix, Daniel Harari, Spandan Madan, Hanspeter Pfister, Tomotake Sasaki, Pawan Sinha
arXiv:2606. 29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of deep learning.
By Peilin Liu, Ding-Xuan Zhou
arXiv:2602. 06065v3 Announce Type: replace-cross Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning.
By Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2606. 14259v1 Announce Type: new Abstract: Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties.
By Chenxiang Zhang, Rustem Islamov, Enea Monzio Compagnoni, Jun Pang, Aurelien Lucchi, Antonio Orvieto
arXiv:2605. 22472v2 Announce Type: replace Abstract: Winner-take-all (WTA) networks constitute a central circuit motif in cortical networks of the brain.
By Julian Gutheil (Graz University of Technology), Simon Hitzginger (Graz University of Technology), Robert Legenstein (Graz University of Technology)