arXiv AI By Mansi Sakarvadia, Aswathy Ajith, Arham Khan, Nathaniel Hudson, Caleb Geniesse, Kyle Chard, Yaoqing Yang, Ian Foster, Michael W. Mahoney

Mitigating Memorization In Language Models

Read the original on arXiv AI →

The paper explores ways to reduce the memorization of training data in language models, testing three regularizer-based, three finetuning-based, and eleven machine unlearning methods—five of which are newly introduced. It introduces TinyMem, a lightweight suite of small models for rapid testing of these mitigation techniques, and shows that unlearning methods, particularly BalancedSubnet, outperform others in removing memorized content while maintaining task performance. The study also finds that regularizer-based approaches are slow and ineffective, while finetuning methods are costly, especially when high accuracy is required.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

Forgetting Only What Matters: Layer-Selective Unlearning toward Robust LLMs

The paper introduces Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer‑selective unlearning framework for large language models. FOM-UL uses a forget‑to‑retain significance score to identify transformer layers that strongly influence the forget set while being insensitive to the retain set, allowing targeted updates that preserve most of the model. Experiments on TOFU, KnowUnDo, and MUSE-style benchmarks show that FOM-UL reduces residual memorization and maintains utility better than several baselines, even after 8‑bit and 4‑bit post‑training quantization, and it also limits recovery of forgotten content in adversarial prompt tests.

By Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou