arXiv AI By Francesco Sovrano, Gabriele Dominici, Marc Langheinrich

Neuron-Anchored Rule Extraction for Large Language Models via Contrastive Hierarchical Ablation

Read the original on arXiv AI →

arXiv:2605. 03058v2 Announce Type: replace-cross Abstract: A central goal of explainable AI is to express large language model (LLM) decision logic symbolically and ground it in internal mechanisms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow
arXiv Computation and Language
Sep 3

GAPS: Dimension-Level Gates for Conditional Activation Steering

The paper introduces GAPS, a dimension‑level gating approach for activation steering in language models. GAPS uses two training‑free gates—a static separability gate based on AUROC and a dynamic posterior gate based on a Gaussian model—to selectively apply steering vectors only to neurons that carry reliable concept information or are currently mis‑activated. Experiments on Gemma‑3 and Qwen‑3 show that GAPS improves or matches the performance of token‑level methods, notably reducing Gemma‑3’s toxicity rate from 6.52% to 0.48% under a fixed capability budget.

By Moghis Fereidouni, Muhammad Umair Haider, Hassan Sajjad, A. B. Siddique