arXiv AI By Amin Omidvar

SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks

Read the original on arXiv AI →

arXiv:2510. 22450v3 Announce Type: replace-cross Abstract: The choice of activation function plays a critical role in neural networks, yet most architectures still rely on fixed, uniform activation functions across all neurons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 8

FlexAct: Why Learn when you can Pick?

arXiv:2601. 06441v2 Announce Type: replace Abstract: Learning activation functions has emerged as a promising direction in deep learning, allowing networks to adapt activation mechanisms to task-specific demands.

By Ramnath Kumar, Kyle Ritscher, Junmin Judy, Lawrence Liu, Cho-Jui Hsieh
arXiv Machine Learning
Sep 22

The Ups and Downs of Backprop Weights

The paper discusses how backpropagation enables deep learning but does not inherently organize parameters for reusable functional components, leading to weight entanglement where overlapping parameter sets hinder independent modification. It introduces weight operators—parameterized modules that can be composed at inference—to address this, proposing a two-stage learning process that first infers operator composition and then updates only the selected operators. Vector Networks (VNs) are presented as an implementation that couples operator selection to local error-driven updates, demonstrating that learned operators can be recombined in unseen ways while keeping updates confined to the relevant parameter sets.

By Giuseppe Chindemi, Benjamin F. Grewe
arXiv Machine Learning
Aug 31

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.

By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong