The paper introduces Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer‑selective unlearning framework for large language models. FOM-UL uses a forget‑to‑retain significance score to identify transformer layers that strongly influence the forget set while being insensitive to the retain set, allowing targeted updates that preserve most of the model. Experiments on TOFU, KnowUnDo, and MUSE-style benchmarks show that FOM-UL reduces residual memorization and maintains utility better than several baselines, even after 8‑bit and 4‑bit post‑training quantization, and it also limits recovery of forgotten content in adversarial prompt tests.
By Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou
arXiv:2607. 02513v1 Announce Type: cross Abstract: LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods.
By Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers
arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.
By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers...
The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.
By Manit Baser, Aditya Nawal, Dinil Mon Divakaran, Mohan Gurusamy
arXiv:2608. 05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs).
By Pawe{\l} Batorski, Przemys{\l}aw Spurek, Paul Swoboda
arXiv:2605. 07482v2 Announce Type: replace Abstract: Machine unlearning for large language models (LLMs) aims to selectively remove memorized content such as private data, copyrighted text, or hazardous knowledge, without costly full retraining.
By Zizhao Hu, Ameya Godbole, Johnny Tian-Zheng Wei, Mohammad Rostami, Jesse Thomason, Robin Jia
The paper introduces Unmerge, an efficient machine unlearning algorithm that treats unlearning as the inverse of task arithmetic. By representing the forget component as a low‑rank basis at each layer, Unmerge optimizes three goals—matching the merged vector, suppressing leakage, and bounding correction size—to limit forget leakage and retain damage. Experiments on ResNet‑50, ViT‑S/16, and Llama‑3.2‑3B show significant performance gains over existing methods while maintaining privacy and feature‑distribution fidelity.
By Haoran Tang, Andrew Tan, Rajiv Khanna
arXiv:2606. 17168v3 Announce Type: replace Abstract: When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary cannot restore it.
By Filip Sondej, Yushi Yang, Adam Mahdi
The paper investigates how machine unlearning for large language models (LLMs) can unintentionally erase related knowledge, even in distant domains. By analyzing the propagation of unlearning effects before any model updates, the authors discover a consistent decay pattern where collateral damage is strongest near the targeted forget set and diminishes with semantic distance but never fully disappears at domain boundaries. They propose a pre-unlearning prediction task—forget-set auditing—to identify potential collateral damage early, finding that interaction features between the forget set and evaluation set are the most predictive signals. This approach offers an early warning system for risky unlearning runs and guides the design of more reliable unlearning procedures.
By Bo Su, Ankit Shah, Thai Le
arXiv:2511. 05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs).
By Yinjie Cheng, Paul Youssef, Christin Seifert, J\"org Schl\"otterer, Zhixue Zhao
arXiv:2606. 03808v1 Announce Type: cross Abstract: We propose PURGE, a machine unlearning algorithm built on a simple but an under-exploited observation: continual learning (CL) and machine unlearning (MU) which are fundamentally dual problems.
By Vedant Jawandhia, Daksh Ahuja, Ghufran Alam Siddiqui, Prashant Trivedi, Yash Sinha, Pratik Narang