arXiv:2510. 22450v3 Announce Type: replace-cross Abstract: The choice of activation function plays a critical role in neural networks, yet most architectures still rely on fixed, uniform activation functions across all neurons.
By Amin Omidvar
The paper discusses how backpropagation enables deep learning but does not inherently organize parameters for reusable functional components, leading to weight entanglement where overlapping parameter sets hinder independent modification. It introduces weight operators—parameterized modules that can be composed at inference—to address this, proposing a two-stage learning process that first infers operator composition and then updates only the selected operators. Vector Networks (VNs) are presented as an implementation that couples operator selection to local error-driven updates, demonstrating that learned operators can be recombined in unseen ways while keeping updates confined to the relevant parameter sets.
By Giuseppe Chindemi, Benjamin F. Grewe
arXiv:2608. 10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain shifts.
By Qiang Wang, Songlin Dong, Shaokun Wang, Jizhou Han, Xiang Song, Chenhao Ding, Yuhang He, Yihong Gong
arXiv:2606. 09928v1 Announce Type: cross Abstract: The Forward-Forward (FF) algorithm offers a biologically inspired alternative to backpropagation by replacing gradient-based credit assignment with local, forward-only objectives.
By Mohammadnavid Ghader, Saeed Reza Kheradpisheh, Bahar Farahani, Mahmood Fazlali
arXiv:2607. 10077v1 Announce Type: new Abstract: Tabular learning is still dominated by gradient-boosted decision trees (GBDTs), while recent deep learning approaches have become increasingly competitive.
By Jiaqi Luo, Shixin Xu
arXiv:2511. 15941v2 Announce Type: replace-cross Abstract: Tabular data underpins decisions across science, industry, and public services.
By David Bonet, Mar\c{c}al Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat, Alexander G. Ioannidis
The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.
By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
Many fine-grained recognition tasks contain hierarchical labels such as order, family and species. Although this supervision should be beneficial, jointly optimising all levels often leads to unstable training because coarse and fine classifiers impose inconsistent gradients on the shared backbone.
arXiv:2604. 11613v4 Announce Type: replace-cross Abstract: Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque.
By Patrick Lutz, Themistoklis Haris, Arjun Chandra, Aditya Gangrade, Venkatesh Saligrama
arXiv:2207. 12877v3 Announce Type: replace Abstract: Motivated by the successes of deep learning, we propose a class of neural network-based discrete choice models, called RUMnets, inspired by the random utility maximization (RUM) framework.
By Ali Aouad, Antoine D\'esir
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
arXiv:2509. 23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent.
By Matt L. Sampson, Peter Melchior