arXiv AI By Greg Kocher, Robert West, Cl\'ement Dumas, Julian Minder

Diff Mining: Logit Differences Reveal Finetuning Objectives

Read the original on arXiv AI →

Diff Mining is a framework that identifies what a finetuned language model has learned by comparing its logits to those of its base model. It extracts per-context logit differences on a reference corpus and aggregates them into an interpretable token set using either a Top‑K frequency method or Non‑negative Matrix Factorization. The approach outperforms existing model‑diffing methods in domain detection and bias identification, and it requires only access to output logits, making it scalable to large models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 17

Combating Data Laundering in LLM Training

arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.

By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu