arXiv:2601. 22594v2 Announce Type: replace-cross Abstract: The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986).
By Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah Schwettmann
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs.
The paper studies GLU-based neurons in large language models by measuring the cosine similarity between each neuron's input and output weight vectors. A strong negative similarity identifies a "weakening neuron," which tends to appear in late layers, activates frequently, and exerts a large influence on model behavior. The authors also find that weakening neurons significantly affect outputs when gate values are negative, contrary to expectations.
By Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui, Hinrich Sch\"utze
arXiv:2507. 00322v2 Announce Type: replace-cross Abstract: Despite remarkable advances in coding capabilities, language models (LMs) still struggle with simple syntactic tasks such as generating balanced parentheses.
By Daking Rai, Samuel Miller, Kevin Moran, Ziyu Yao
arXiv:2606. 19317v1 Announce Type: cross Abstract: A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions.
By Amiri Hayes, Belinda Li, Jacob Andreas