arXiv AI

Diff Mining: Logit Differences Reveal Finetuning Objectives

Diff Mining is a framework that identifies what a finetuned language model has learned by comparing its logits to those of its base model. It extracts per-context logit differences on a reference corpus and aggregates them into an interpretable token set using either a Top‑K frequency method or Non‑negative Matrix Factorization. The approach outperforms existing model‑diffing methods in domain detection and bias identification, and it requires only access to output logits, making it scalable to large models.

arXiv AI
Jun 17

Combating Data Laundering in LLM Training

arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.

By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu
arXiv Machine Learning
Sep 21

Exemplar Partitioning for Mechanistic Interpretability

The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.

By Jessica Rumbelow
arXiv Machine Learning
Aug 19

J-Miner: Recovering Executable Decision Knowledge from Language-Model Classifiers

J-Miner extracts executable decision knowledge from fine‑tuned language‑model classifiers by aggregating internal signals aligned with named concepts and learning rules over them. The mined rules reproduce up to 98.3% of the original classifier’s decisions and outperform compact rules based on input words by 6.0–29.5 percentage points in behavioral fidelity. Moreover, these rules can be transferred to lightweight student models that use only about 1/24 of the parameters yet retain 99.8% of the source classifier’s accuracy.

By Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang