arXiv Machine Learning

A Prior-Aware Metric for Efficiently Distinguishing Memorization from Generalization in Large Language Models

arXiv:2602. 18733v2 Announce Type: replace Abstract: Training data leakage from Large Language Models (LLMs) raises serious concerns related to privacy, security, and copyright compliance.

arXiv Machine Learning
Jul 15

Extractable Memorization From First Principles

arXiv:2607. 12649v1 Announce Type: new Abstract: Recent work on extractable memorization in LLMs suffers from two contrasting validity problems.

By A. Feder Cooper, Marika Swanberg, Jamie Hayes, Lea Duesterwald, Christopher De Sa, Daniel E. Ho, Mark A. Lemley, Percy Liang
arXiv AI
3d ago

Mitigating Memorization In Language Models

The paper explores ways to reduce the memorization of training data in language models, testing three regularizer-based, three finetuning-based, and eleven machine unlearning methods—five of which are newly introduced. It introduces TinyMem, a lightweight suite of small models for rapid testing of these mitigation techniques, and shows that unlearning methods, particularly BalancedSubnet, outperform others in removing memorized content while maintaining task performance. The study also finds that regularizer-based approaches are slow and ineffective, while finetuning methods are costly, especially when high accuracy is required.

By Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Nathaniel Hudson, Caleb Geniesse, Kyle Chard, Yaoqing Yang, Ian Foster, Michael W. Mahoney
arXiv Machine Learning
Jul 22

Estimating near-verbatim extraction risk in language models with decoding-constrained beam search

arXiv:2603. 24917v2 Announce Type: replace-cross Abstract: Recent work shows that standard greedy-decoding extraction methods for quantifying memorization in LLMs miss how extraction risk varies across sequences.

By A. Feder Cooper, Mark A. Lemley, Christopher De Sa, Lea Duesterwald, Allison Casasola, Jamie Hayes, Katherine Lee, Daniel E. Ho, Percy Liang
arXiv AI
Sep 10

Alignment Whack-a-Mole : Finetuning Activates Verbatim Recall of Copyrighted Books in Large Language Models

The paper demonstrates that fine‑tuning large language models on a single author’s works can trigger the models to reproduce large verbatim excerpts from copyrighted books, even when prompted only with semantic descriptions. Experiments on GPT‑4o, Gemini‑2.5‑Pro, and DeepSeek‑V3.1 show up to 85‑90% recall of held‑out books, with spans exceeding 460 words, and this effect generalizes across authors and model providers. The findings suggest that fine‑tuning reactivates latent memorization from pre‑training, revealing a widespread vulnerability in industry models.

By Xinyue Liu, Niloofar Mireshghallah, Jane C. Ginsburg, Tuhin Chakrabarty
arXiv Machine Learning
Aug 19

Quantifying Memorization and Privacy Risks in Genomic Language Models

The paper introduces a comprehensive privacy evaluation framework for genomic language models (GLMs) that quantifies memorization risks using perplexity-based detection, canary sequence extraction, and membership inference. By planting canary sequences at different repetition rates in synthetic and real datasets, the authors systematically assess how repetition, model capacity, and training dynamics affect memorization across various GLM architectures. The study demonstrates that GLMs do memorize training data to varying degrees and that no single attack method fully captures this risk, highlighting the necessity of multi-vector privacy auditing for genomic AI systems.

By Alexander Nemecek, Wenbiao Li, Xiaoqian Jiang, Jaideep Vaidya, Erman Ayday
arXiv Machine Learning
Jul 27

DCS: A Unified Conditional Sensitivity Framework for Cross-Modal Copyright Infringement Detection

arXiv:2607. 22035v1 Announce Type: new Abstract: Currently, most foundation models can reproduce or strongly depend on copyrighted training content, but output similarity alone is insufficient for infringement detection, because similar outputs may also arise from public-domain concepts, common stylistic conventions, or ordinary statistical generalization.

By Xiafeng Man