CALIBURN is a new approach to large language model (LLM) unlearning that measures a model’s confidence in undesirable knowledge and uses this measure to fine‑tune unlearning gradient updates. By doing so, it offers more precise control over what is forgotten while better preserving the model’s overall utility. Experiments on benchmarks such as MUSE and WMDP show that CALIBURN outperforms existing methods in balancing knowledge removal with utility retention.
By Zhengbang Yang, Yisheng Zhong, Junyuan Hong, Zhuangdi Zhu
arXiv:2609.27355v1 Announce Type: cross
Abstract: Unlearning ensures LLM compliance by removing the influence of private or copyrighted training data. However, since LLM models typically undergo post...
By Jialu Wang, Jianing Deng, Shuqing Luo, Yuanzhe Li, Dongwei Wang, Jingtong Hu, Huanrui Yang, Song Wang, Tianlong Chen
The paper introduces Forgetting Only What Matters via Unlearning Layers (FOM-UL), a layer‑selective unlearning framework for large language models. FOM-UL uses a forget‑to‑retain significance score to identify transformer layers that strongly influence the forget set while being insensitive to the retain set, allowing targeted updates that preserve most of the model. Experiments on TOFU, KnowUnDo, and MUSE-style benchmarks show that FOM-UL reduces residual memorization and maintains utility better than several baselines, even after 8‑bit and 4‑bit post‑training quantization, and it also limits recovery of forgotten content in adversarial prompt tests.
By Ravi Ranjan, Olivera Kotevska, Agoritsa Polyzou
The paper investigates whether a model that has undergone class unlearning can still recover forgotten classes without access to original data. It introduces a white‑box audit method that generates synthetic probes in representation space, filters them by confidence, and relabels boundary‑adjacent probes as the forgotten class. The authors define a Relearning Score to quantify recovery while preserving retain performance, and demonstrate that several unlearning techniques on CIFAR‑10, CIFAR‑100, and TinyImageNet can be fully recovered in a source‑free setting, sometimes even outperforming a retrained reference.
By Zahra Dehghani, Pablo Piantanida, Mohammadhadi Shateri
Large Language Models (LLMs) can memorize and reproduce sensitive, copyrighted, or otherwise undesirable training content, creating privacy, safety, and regulatory concerns. Machine unlearning offers...
The paper identifies a problem in large language model (LLM) unlearning called forget‑set misalignment, where the set of data to be forgotten does not match what the model has actually memorized. Two failure modes are described: Under Unlearning, where memorized information is omitted from the forget set, and Out‑of‑Knowledge Unlearning, where the algorithm attempts to forget knowledge the model never learned, harming performance. The authors propose CONfs, a data‑blind framework that constructs model‑aligned forget sets by eliciting the model’s memorized knowledge, and demonstrate that it achieves near‑gold standard forgetting while preserving utility better than other data‑blind methods.
By Miso Kim, Georu Lee, Seungwon Jeong, Woojin Lee
The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.
By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv:2608.28361v1 Announce Type: cross
Abstract: Machine Unlearning methods for Large Language Models typically assume pre-specified forget and retain sets. In realistic settings, however, requests...
By Praveen Bushipaka, Andrea D'Angelo, Lucia Passaro, Tommaso Cucinotta
Machine unlearning for large language models (LLMs) often assumes that a pre-defined forget set matches what the model has memorized, but this frequently breaks in realistic privacy settings where the...
arXiv:2507. 07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores.
By Jaeheun Jung, Bosung Jung, Suhyun Bae, Donghun Lee
arXiv:2606. 17168v3 Announce Type: replace Abstract: When LLM weights are open or fine-tuning is available through an API, suppressing hazardous knowledge and tendencies is not enough: removal has to be deep enough that an adversary cannot restore it.
By Filip Sondej, Yushi Yang, Adam Mahdi
The paper introduces Conflict-Aware Unlearning (CAU), a schema‑aware method designed to reduce forget‑retain interference in tabular data. CAU relaxes preservation constraints on retained rows that conflict most with the forget set, improving alignment with a retraining oracle while preserving predictive utility. Experiments on clinical and non‑medical tabular tasks demonstrate CAU’s effectiveness for both sample‑level and feature‑level unlearning.
By Zijie Liu, Jinhao Duan, Bingqi Shang, Xinming An, Sijia Liu, Tianlong Chen