arXiv Machine Learning

IGLU: The Integrated Gaussian Linear Unit Activation Function

arXiv:2603. 06861v2 Announce Type: replace Abstract: Activation functions are fundamental to deep neural networks, governing gradient flow, optimization stability, and representational capacity.

arXiv Machine Learning
Sep 11

SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations

SG-Blend introduces a per‑layer adaptive activation that interpolates between a bias‑corrected, parametric Swish variant (SSwish) and GELU, using a learnable blend coefficient, sharpness, and zero‑centering bias. The method adds only three scalars per feed‑forward block and, on BERT‑style IMDB classification, matches peak accuracy while reducing seed‑to‑seed variance by 42 %. It also achieves the lowest validation perplexity on WikiText103 and generalizes to computer vision and other domains.

By Gaurav Sarkar, Syed Affan Daimi, Jay Gala, Subarna Tripathi
arXiv Machine Learning
Aug 31

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.

By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
arXiv Machine Learning
Jun 3

Dynamic Short Convolutions Improve Transformers

arXiv:2606. 03825v1 Announce Type: new Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization.

By Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried
arXiv Machine Learning
Jul 15

Inference-Time Machine Unlearning via Gated Activation Redirection

arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.

By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u