arXiv Machine Learning
4d ago

Reference-Guided Machine Unlearning

Reference-Guided Machine Unlearning (ReGUn) is a vision unlearning framework that prioritizes distributional indistinguishability over degradation-based heuristics. It uses disjoint held-out data to create a class-conditioned reference distribution for distillation, guiding forget samples toward non-member behavior without explicitly degrading predictions. Experiments across various architectures and datasets show that ReGUn achieves a competitive forgetting–utility trade-off and closely matches retrain-like membership inference behavior.

By Jonas Mirlach, Sonia Laguna, Julia E. Vogt
arXiv Machine Learning
Sep 3

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. Using Selective Gradient Masking, the authors train six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge, then apply three standard unlearning methods to each. Results show that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost at the same forgetting level, providing direct evidence that entanglement contributes to collateral damage in unlearning.

By Ev\v{z}en Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
Hugging Face Trending Papers
Sep 2

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. By training six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge and applying three unlearning methods, the authors find that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost. This controlled experiment provides direct evidence that entanglement contributes to collateral damage during unlearning, supporting a long‑standing hypothesis in interpretability research.