OpenAI Blog

Language models can explain neurons in language models

Read the original on OpenAI Blog →

We use GPT-4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. We release a dataset of these (imperfect) explanations and scores for every neuron in GPT-2.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

Hugging Face Trending Papers
Jun 17

Explaining Attention with Program Synthesis

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs.

arXiv Machine Learning
Sep 17

Weakening Neurons: An Input-Output Functionality in Transformers with Outsize Influence

The paper studies GLU-based neurons in large language models by measuring the cosine similarity between each neuron's input and output weight vectors. A strong negative similarity identifies a "weakening neuron," which tends to appear in late layers, activates frequently, and exerts a large influence on model behavior. The authors also find that weakening neurons significantly affect outputs when gate values are negative, contrary to expectations.

By Sebastian Gerstner, Hilal AlQuabeh, Kentaro Inui, Hinrich Sch\"utze