arXiv Machine Learning

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. Using Selective Gradient Masking, the authors train six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge, then apply three standard unlearning methods to each. Results show that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost at the same forgetting level, providing direct evidence that entanglement contributes to collateral damage in unlearning.

Hugging Face Trending Papers
Sep 2

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. By training six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge and applying three unlearning methods, the authors find that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost. This controlled experiment provides direct evidence that entanglement contributes to collateral damage during unlearning, supporting a long‑standing hypothesis in interpretability research.

arXiv AI
Aug 28

What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs

The paper investigates how the structural entanglement of facts within a large language model’s knowledge base influences whether those facts leak after unlearning. Using two unlearning algorithms (WHP and GA+KL) across fictional and real-world datasets and multiple model sizes, the authors find that highly entangled facts are more likely to be recalled before unlearning, but the relationship changes—WHP weakens it while GA+KL reverses it. By directly manipulating entanglement scores and observing corresponding recall changes, they demonstrate a causal link and develop a predictive tool to audit prompts for potential leakage.

By Aakriti Shah, Yifan Hu, Thai Le
arXiv Machine Learning
Jun 15

Natively Unlearnable Large Language Models

arXiv:2606. 13873v1 Announce Type: new Abstract: Unlearning aims to remove the influence of specific training data sources, but this has proved challenging because the contributions of different sources are entangled within the model.

By Gaurav R. Ghosal, Pratyush Maini, Aditi Raghunathan
arXiv Computation and Language
Sep 23

PreUnlearn: Auditing Collateral Knowledge Damage Before Large Language Model Unlearning

The paper investigates how machine unlearning for large language models (LLMs) can unintentionally erase related knowledge, even in distant domains. By analyzing the propagation of unlearning effects before any model updates, the authors discover a consistent decay pattern where collateral damage is strongest near the targeted forget set and diminishes with semantic distance but never fully disappears at domain boundaries. They propose a pre-unlearning prediction task—forget-set auditing—to identify potential collateral damage early, finding that interaction features between the forget set and evaluation set are the most predictive signals. This approach offers an early warning system for risky unlearning runs and guides the design of more reliable unlearning procedures.

By Bo Su, Ankit Shah, Thai Le