arXiv Computation and Language

Forging LLM Authorship Fingerprints with Targeted Rewriting

arXiv AI
Jun 17

Combating Data Laundering in LLM Training

arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.

By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu
arXiv Machine Learning
1d ago

Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution

The study investigates whether provenance information can reliably identify the source of synthetic text and whether this identification improves the selection of training data. Using financial‑risk text, the authors achieve 98.7% accuracy in attributing original generated passages, but accuracy drops to 53.1% after paraphrasing and 29.0% after style rewriting. They compare two selection strategies—one based on source provenance and another on a reference model score—across three rounds of generation and retraining, finding that the two methods choose different examples but do not produce a consistent difference in model degradation. The results suggest that source attribution and useful data selection are distinct challenges, and neither provenance nor the tested proxy suffices to guarantee stable recursive training behavior.

By Joss Armstrong
arXiv AI
Sep 10

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

The paper demonstrates that fine‑tuning large language models on a single author’s works can trigger the models to reproduce large verbatim excerpts from copyrighted books, even when prompted only with semantic descriptions. Experiments on GPT‑4o, Gemini‑2.5‑Pro, and DeepSeek‑V3.1 show up to 85‑90% recall of held‑out books, with spans exceeding 460 words, and this effect generalizes across authors and model providers. The findings suggest that fine‑tuning reactivates latent memorization from pre‑training, revealing a widespread vulnerability in industry models.

By Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty
arXiv AI
Aug 11

Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification

arXiv:2608. 08195v1 Announce Type: cross Abstract: Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignment.

By Yutong Wu, Xiaofan Bai, Shixin Li, Pingyi Hu, Ziqi Zhou, Zilong Wang, Xiaojing Ma, Songfeng Lu, Yuhong Li, Jin Xuan, Yi Wang, Dongmei Zhang, Bin Benjamin Zhu