arXiv AI

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

arXiv:2607. 13037v1 Announce Type: new Abstract: When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author.

arXiv AI
Aug 24

Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning

The paper introduces the task of Scientific Claim Unlearning and presents a new benchmark, SciUnlearn, to evaluate it. It highlights that language models trained on static scientific corpora risk disseminating outdated or retracted claims as scientific knowledge evolves. Current machine unlearning methods fail to effectively remove claim-level knowledge, often only suppressing it superficially, underscoring the need for specialized techniques for structured knowledge removal.

By Snigdha Paul, Manasi Patwardhan, Arman Cohan
arXiv AI
Aug 26

Provenance Guided Incremental Learning Under Evolving Concept Definitions

The paper introduces a provenance-guided incremental learning framework that handles rule-induced concept shift, where target definitions are explicitly revised and previously stored instances receive new semantic labels. By compiling concept changes into structured rule deltas, tracing affected records through historical provenance, and selectively re-evaluating only a localized candidate region, the method automatically relabels executable revisions, manages ambiguous cases with selective supervision, and repairs predictors incrementally. Evaluation on the RuleShift-Bench benchmark—covering financial, demographic, cybersecurity, and graph-structured data—shows 92.3% accuracy and 90.2% Macro‑F1, reprocessing only 14.7% of the historical collection and achieving an average update latency of 179 s versus 993 s for full relabeling and retraining.

By Ismail Lamaakal
arXiv AI
Aug 28

Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds

The paper introduces GRAPHSU, a graph‑guided selective unlearning method for language models that expands deletion beyond explicitly identified forget seeds. By constructing a weighted support‑route graph and propagating deletion pressure, GRAPHSU applies graded forgetting to high‑risk neighboring examples. Experiments on the TOFU and PISTOL benchmarks with GPT‑2 Medium and Llama‑3.2‑3B‑Instruct show that GRAPHSU achieves the lowest utility‑feasible soft leakage, reducing leakage by up to 49.5 percentage points compared to a seed‑only baseline.

By Waqas Khan, Tabinda Sarwar, Jingyue Cong, Xun Yi, Estrid He
arXiv AI
Aug 12

Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data

arXiv:2608. 11022v1 Announce Type: cross Abstract: Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use.

By Nicola Giuseppe Marchioro, Gabriele Padovani, Amal Gueroudji, Rafael Ferreira da Silva, Wesley Brewer, Valentine Anantharaj, Sandro Fiore, Renan Souza
arXiv AI
Jun 17

Combating Data Laundering in LLM Training

arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.

By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu