arXiv Machine Learning

Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing

arXiv Computation and Language
Aug 25

Can LLMs Truly Forget? Revealing Unlearning Gaps Through Adversarial Evaluation

arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whet...

By Ayush Gupta, Hima Varshini Surisetty, Sreevidya Bollineni, Varad Ingale, Tuhina Tripathi, Abhishek Lalwani, Somya Chatterjee, Sadid Hasan
arXiv AI
Sep 1

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

The paper introduces FAB, an attack that uses meta‑learning to embed dormant adversarial behaviors into large language models (LLMs). These behaviors remain inactive until the model is finetuned by downstream users, at which point the model can exhibit unwanted actions such as unsolicited advertising, jailbreakability, or over‑refusal. FAB is shown to be effective across multiple LLMs and resilient to various finetuning settings.

By Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev
arXiv AI
3d ago

Faithful Dual-constrained Erasure for Robust LLM Safety Alignment

The paper introduces FDCU, a dual‑constrained subspace projection framework designed to improve machine unlearning for large language models. FDCU limits parameter updates with a dual‑masking rule that preserves general knowledge via Fisher Information while preventing the activation of spurious suppressors through the Principle of Minimal Functional Intervention. Experiments show that FDCU achieves state‑of‑the‑art robustness against retraining attacks while maintaining near‑lossless general utility, thereby ensuring durable safety for LLMs.

By Jiaqing Li, Shide Zhou, Zhibo Zhang, Yuxi Li, Tianlong Yu, Kailong Wang
arXiv Machine Learning
Jul 15

Inference-Time Machine Unlearning via Gated Activation Redirection

arXiv:2605. 12765v3 Announce Type: replace Abstract: Large Language Models memorize vast amounts of training data, raising concerns regarding privacy, copyright infringement, and safety.

By Vin\'icius Conte Turani, Ot\'avio Parraga, Jo\~ao Vitor Boer Abitante, Kristen K. Arguello, Joana Pasquali, Ramiro N. Barros, Flavio du Pin Calmon, Christian Mattjie, Rodrigo C. Barros, Lucas S. Kupssinsk\"u
arXiv AI
Sep 25

The Tokens Remember: When Tokenization Bypasses Knowledge Editing and Unlearning

The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.

By Manit Baser, Aditya Nawal, Dinil Mon Divakaran, Mohan Gurusamy