Rethinking Neural Nonlinearity as Gating
arXiv:2607. 03148v1 Announce Type: cross Abstract: Activation functions are considered an essential primitive for neural nonlinearity, i.
arXiv:2607. 03664v1 Announce Type: new Abstract: The Gaussian Error Linear Unit is usually motivated as the expected output of an input-dependent stochastic Bernoulli gate.
arXiv:2607. 03148v1 Announce Type: cross Abstract: Activation functions are considered an essential primitive for neural nonlinearity, i.
The paper introduces GAPS, a dimension‑level gating approach for activation steering in language models. GAPS uses two training‑free gates—a static separability gate based on AUROC and a dynamic posterior gate based on a Gaussian model—to selectively apply steering vectors only to neurons that carry reliable concept information or are currently mis‑activated. Experiments on Gemma‑3 and Qwen‑3 show that GAPS improves or matches the performance of token‑level methods, notably reducing Gemma‑3’s toxicity rate from 6.52% to 0.48% under a fixed capability budget.
arXiv:2603. 06861v2 Announce Type: replace Abstract: Activation functions are fundamental to deep neural networks, governing gradient flow, optimization stability, and representational capacity.
arXiv:2606.28444v2 Announce Type: replace-cross Abstract: Classical universal approximation theorems (UAT) establish the expressive power of sigmoidal multilayer perceptrons, but they do not specify...
arXiv:2606. 20292v1 Announce Type: new Abstract: The use of neural networks (NNs) is rapidly increasing, including in safety- and security-critical domains.
arXiv:2602.17493v2 Announce Type: replace-cross Abstract: We develop a method for training neural networks on Boolean data in which the values at all nodes are strictly $\pm 1$, and the resulting mod...
arXiv:2510. 22450v3 Announce Type: replace-cross Abstract: The choice of activation function plays a critical role in neural networks, yet most architectures still rely on fixed, uniform activation functions across all neurons.
arXiv:2602. 22352v2 Announce Type: replace-cross Abstract: With the continuous growth of neural network scales, low-precision quantization is widely used in edge accelerators.
arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
arXiv:2606. 28444v1 Announce Type: cross Abstract: Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe how initial weights should encode the geometry of a data distribution.
arXiv:2607. 09399v1 Announce Type: cross Abstract: We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and lookup table networks (LUTNs).
arXiv:2607. 15525v1 Announce Type: cross Abstract: Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interface for interpretation and a possible alternative to transformer feed-forward networks.