PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning
Read the original on arXiv Computation and Language →The paper investigates how machine unlearning for large language models (LLMs) can unintentionally erase related knowledge, even in distant domains. By analyzing the propagation of unlearning effects before any model updates, the authors discover a consistent decay pattern where collateral damage is strongest near the targeted forget set and diminishes with semantic distance but never fully disappears at domain boundaries. They propose a pre-unlearning prediction task—forget-set auditing—to identify potential collateral damage early, finding that interaction features between the forget set and evaluation set are the most predictive signals. This approach offers an early warning system for risky unlearning runs and guides the design of more reliable unlearning procedures.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.