arXiv AI By Xujia Li, Dan Li, Jian Lou, Wenjie Feng

Signal-Guided Optimization for Machine Unlearning

Read the original on arXiv AI →

arXiv:2607. 11975v1 Announce Type: cross Abstract: Current machine unlearning methods predominantly rely on global, coarse-grained intervention strategies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.

By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv AI
Sep 2

Confess What You Know: Forget-Set Misalignment with Model Knowledge in LLM Unlearning

The paper identifies a problem in large language model (LLM) unlearning called forget‑set misalignment, where the set of data to be forgotten does not match what the model has actually memorized. Two failure modes are described: Under Unlearning, where memorized information is omitted from the forget set, and Out‑of‑Knowledge Unlearning, where the algorithm attempts to forget knowledge the model never learned, harming performance. The authors propose CONfs, a data‑blind framework that constructs model‑aligned forget sets by eliciting the model’s memorized knowledge, and demonstrate that it achieves near‑gold standard forgetting while preserving utility better than other data‑blind methods.

By Miso Kim, Georu Lee, Seungwon Jeong, Woojin Lee
arXiv Machine Learning
4d ago

Reference-Guided Machine Unlearning

Reference-Guided Machine Unlearning (ReGUn) is a vision unlearning framework that prioritizes distributional indistinguishability over degradation-based heuristics. It uses disjoint held-out data to create a class-conditioned reference distribution for distillation, guiding forget samples toward non-member behavior without explicitly degrading predictions. Experiments across various architectures and datasets show that ReGUn achieves a competitive forgetting–utility trade-off and closely matches retrain-like membership inference behavior.

By Jonas Mirlach, Sonia Laguna, Julia E. Vogt