Contrastive Representation Shaping for LLM Unlearning
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2508.20443v3 Announce Type: replace Abstract: Large language models (LLMs) are trained on massive datasets that may include private or copyrighted content. Due to growing privacy and ownership...
arXiv:2507. 07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores.
Reference-Guided Machine Unlearning (ReGUn) is a vision unlearning framework that prioritizes distributional indistinguishability over degradation-based heuristics. It uses disjoint held-out data to create a class-conditioned reference distribution for distillation, guiding forget samples toward non-member behavior without explicitly degrading predictions. Experiments across various architectures and datasets show that ReGUn achieves a competitive forgetting–utility trade-off and closely matches retrain-like membership inference behavior.
The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. Using Selective Gradient Masking, the authors train six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge, then apply three standard unlearning methods to each. Results show that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost at the same forgetting level, providing direct evidence that entanglement contributes to collateral damage in unlearning.
The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. By training six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge and applying three unlearning methods, the authors find that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost. This controlled experiment provides direct evidence that entanglement contributes to collateral damage during unlearning, supporting a long‑standing hypothesis in interpretability research.
arXiv:2609.05966v1 Announce Type: new Abstract: Existing unlearning approaches typically rely on post hoc weight adaptation or distillation, leading to duplicated memory costs, degraded generalizatio...