arXiv:2606. 02765v1 Announce Type: cross Abstract: Model dimension ($d_{model}$) is a fundamental hyperparameter in transformer language models, yet its role in setting the geometric limits of feature representation remains under-explored.
By Alexander Guha
arXiv:2606. 02385v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control.
By William Dorrell
arXiv:2512.04696v3 Announce Type: replace
Abstract: We develop a flexible feature selection framework based on deep neural networks that approximately controls the false discovery rate (FDR), a measu...
By Kazuma Sawaya
arXiv:2607. 05546v1 Announce Type: cross Abstract: We develop a unified function space theory of deep fully connected neural networks.
By Julia Nakhleh, Robert D. Nowak
arXiv:2606. 09725v1 Announce Type: new Abstract: Disentanglement, the separation of factors of variation in data using neural networks, remains a long-standing challenge in machine learning.
By Jhonny J. Velasquez Olivera, Christo K. Thomas, Walid Saad
arXiv:2512. 24780v2 Announce Type: replace Abstract: Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking.
By Alan Oursland
The paper investigates the expressive power of multimodal contrastive learning architectures by treating them as parameterized families of joint density estimators. It shows that the classic two‑tower CLIP model is a universal approximator for two modalities, while a common extension that sums pairwise similarities fails to approximate arbitrary joint distributions when three or more modalities are involved, though it can match all pairwise conditionals. To address this limitation, the authors introduce Hadamard‑CLIP, which adds a single learned weight vector to restore universal approximation for any number of modalities while retaining CLIP’s efficient retrieval capabilities.
By Andrew Stuart, Florian Wolf
arXiv:2605. 22472v2 Announce Type: replace Abstract: Winner-take-all (WTA) networks constitute a central circuit motif in cortical networks of the brain.
By Julian Gutheil (Graz University of Technology), Simon Hitzginger (Graz University of Technology), Robert Legenstein (Graz University of Technology)
arXiv:2606. 19538v1 Announce Type: new Abstract: Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and content-dependent pairwise interaction -- and have remained mathematically distinct since their inception.
By Ashim Dhor, Rasel Mondal, Pin Yu Chen
arXiv:2606. 14954v1 Announce Type: cross Abstract: We develop a general framework for analyzing representation costs of parametric data-fitting methods through their parameter-space regularizers.
By Greg Ongie, Rahul Parhi
arXiv:2606. 21497v2 Announce Type: replace-cross Abstract: Modern deep neural networks are trained using error backpropagation, which requires sequential forward and backward computations across network layers.
By Neeraj Mohan Sushma, Aditya Nagarsekar, Cabrel Teguemne Fokam, Robin Schiewer, Amit Kumar Pal, Anand Subramoney, David Kappel
arXiv:2606. 00130v2 Announce Type: replace-cross Abstract: Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters.
By Andrzej Cichocki, Michal Wietczak