The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. Using Selective Gradient Masking, the authors train six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge, then apply three standard unlearning methods to each. Results show that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost at the same forgetting level, providing direct evidence that entanglement contributes to collateral damage in unlearning.
By Ev\v{z}en Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
The paper investigates how the structural entanglement of facts within a large language model’s knowledge base influences whether those facts leak after unlearning. Using two unlearning algorithms (WHP and GA+KL) across fictional and real-world datasets and multiple model sizes, the authors find that highly entangled facts are more likely to be recalled before unlearning, but the relationship changes—WHP weakens it while GA+KL reverses it. By directly manipulating entanglement scores and observing corresponding recall changes, they demonstrate a causal link and develop a predictive tool to audit prompts for potential leakage.
By Aakriti Shah, Yifan Hu, Thai Le
arXiv:2601.22028v2 Announce Type: replace
Abstract: Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives...
By Haoran Tang, Rajiv Khanna
arXiv:2508.20443v3 Announce Type: replace
Abstract: Large language models (LLMs) are trained on massive datasets that may include private or copyrighted content. Due to growing privacy and ownership...
By Zhihao Liu, Jian Lou, Yuke Hu, Xiaochen Li, Yitian Chen, Tailun Chen, Zhizhen Qin, Kui Ren, Zhan Qin
arXiv:2507. 07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores.
By Jaeheun Jung, Bosung Jung, Suhyun Bae, Donghun Lee
arXiv:2606. 05644v1 Announce Type: new Abstract: When retrieved evidence contradicts parametric memory, language models frequently ignore context and default to memorized priors -- a failure that undermines the core purpose of retrieval augmentation.
By Zhe Yu, Wenpeng Xing, Tiancheng Zhao, Mohan Li, Changting Lin, Meng Han