arXiv Machine Learning By Mingi Kang, Zai Yang, Jeova Farias Sales Rocha Neto

IGLU: The Integrated Gaussian Linear Unit Activation Function

Read the original on arXiv Machine Learning →

arXiv:2603. 06861v2 Announce Type: replace Abstract: Activation functions are fundamental to deep neural networks, governing gradient flow, optimization stability, and representational capacity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

SG-Blend: Learning an Interpolation Between Improved Swish and GELU for Robust Neural Representations

SG-Blend introduces a per‑layer adaptive activation that interpolates between a bias‑corrected, parametric Swish variant (SSwish) and GELU, using a learnable blend coefficient, sharpness, and zero‑centering bias. The method adds only three scalars per feed‑forward block and, on BERT‑style IMDB classification, matches peak accuracy while reducing seed‑to‑seed variance by 42 %. It also achieves the lowest validation perplexity on WikiText103 and generalizes to computer vision and other domains.

By Gaurav Sarkar, Syed Affan Daimi, Jay Gala, Subarna Tripathi
arXiv Machine Learning
Aug 31

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

The paper introduces Mixture of Activations (MoA), a token‑adaptive feedforward network design that mixes multiple activation functions using lightweight gates while sharing linear projections. It also presents learnable activations (LA) as an input‑independent variant. The authors theoretically prove that MoA strictly surpasses both fixed‑activation FFNs and LA in expressive power, and empirically demonstrate that MoA achieves lower loss and better scaling on dense and MoE language models from 0.12 B to 2 B parameters with minimal overhead.

By Mingze Wang, Jinbo Wang, Yikuan Xia, Kai Shen, Shu Zhong
arXiv Machine Learning
Jun 3

Dynamic Short Convolutions Improve Transformers

arXiv:2606. 03825v1 Announce Type: new Abstract: Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization.

By Oliver Sieberling, Bharat Runwal, Rameswar Panda, Yoon Kim
arXiv AI
Jun 15

Gefen: Optimized Stochastic Optimizer

arXiv:2606. 13894v1 Announce Type: cross Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory.

By Nadav Benedek, Tomer Koren, Ohad Fried