Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The paper studies GLU-based neurons in large language models by measuring the cosine similarity between each neuron's input and output weight vectors. A strong negative similarity identifies a "weakening neuron," which tends to appear in late layers, activates frequently, and exerts a large influence on model behavior. The authors also find that weakening neurons significantly affect outputs when gate values are negative, contrary to expectations.
arXiv:2608. 03921v2 Announce Type: replace Abstract: This paper offers a new interpretation of the Transformer during inference.
arXiv:2608. 03921v1 Announce Type: new Abstract: This paper offers a new interpretation of the Transformer during inference.
arXiv:2410. 24050v3 Announce Type: replace Abstract: Large-scale pretraining of transformers has been central to the success of foundation models.
The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.
arXiv:2607. 02964v1 Announce Type: cross Abstract: A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does.