arXiv AI By Amit LeVi, Elad David, Max Fomin

Unsupervised Features Mining via Activation Geometry

Read the original on arXiv AI →

arXiv:2607. 04222v1 Announce Type: new Abstract: Interpretability methods aim to reveal the features represented inside large language models (LLMs).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow
arXiv Machine Learning
Sep 22

Characterizing Model-Native Skills

arXiv:2604.17614v2 Announce Type: replace-cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterization...

By Feiyang Kang, Mahavir Dabas, Myeongseob Ko, Ruoxi Jia