Hugging Face Trending Papers

GAPS: Dimension-Level Gates for Conditional Activation Steering

arXiv Computation and Language
Sep 3

GAPS: Dimension-Level Gates for Conditional Activation Steering

The paper introduces GAPS, a dimension‑level gating approach for activation steering in language models. GAPS uses two training‑free gates—a static separability gate based on AUROC and a dynamic posterior gate based on a Gaussian model—to selectively apply steering vectors only to neurons that carry reliable concept information or are currently mis‑activated. Experiments on Gemma‑3 and Qwen‑3 show that GAPS improves or matches the performance of token‑level methods, notably reducing Gemma‑3’s toxicity rate from 6.52% to 0.48% under a fixed capability budget.

By Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad, A. B. Siddique
Hugging Face Trending Papers
Sep 24

Grammatical "grandmother neurons" are rare in LLMs

The paper introduces a probe‑free method called the Neuron Separability Index (NSI) to assess how individual neurons in large language models distinguish grammatical from ungrammatical sentences using linguistic minimal pairs. Across 68 linguistic paradigms and seven model checkpoints, the study finds that while raw separability for morphology and syntax peaks early, single‑unit selectivity is sparse and weak, with rare strongly selective "grandmother neurons." Moreover, the research shows a dissociation between whole‑vector linear separability, single‑neuron selectivity, and behavioral competence, and demonstrates that targeted ablations can further separate activation selectivity from causal reliance.

arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow