arXiv AI By Aakriti Shah, Yifan Hu, Thai Le

What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs

Read the original on arXiv AI →

The paper investigates how the structural entanglement of facts within a large language model’s knowledge base influences whether those facts leak after unlearning. Using two unlearning algorithms (WHP and GA+KL) across fictional and real-world datasets and multiple model sizes, the authors find that highly entangled facts are more likely to be recalled before unlearning, but the relationship changes—WHP weakens it while GA+KL reverses it. By directly manipulating entanglement scores and observing corresponding recall changes, they demonstrate a causal link and develop a predictive tool to audit prompts for potential leakage.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. Using Selective Gradient Masking, the authors train six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge, then apply three standard unlearning methods to each. Results show that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost at the same forgetting level, providing direct evidence that entanglement contributes to collateral damage in unlearning.

By Ev\v{z}en Wybitul, Tim G. J. Rudner, Christian Schroeder de Witt
Hugging Face Trending Papers
Sep 2

Entangled Representations Amplify Collateral Damage in Unlearning

The paper investigates whether representational entanglement—shared structure between knowledge domains—impedes unlearning in neural networks. By training six 254M‑parameter language models with varying degrees of disentanglement between biology and non‑biology knowledge and applying three unlearning methods, the authors find that more disentangled models consistently achieve better retain‑forget trade‑offs, with up to four‑fold lower retain cost. This controlled experiment provides direct evidence that entanglement contributes to collateral damage during unlearning, supporting a long‑standing hypothesis in interpretability research.

arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin
Hugging Face Trending Papers
Jul 20

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer. We study where this susceptibility, spanning sycophancy and related cue-induced biases, lives inside the model.