Preemptive LLM Unlearning against Forbidden Capability Acquisition via Gradient Sealing
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2512. 14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications.
arXiv:2607. 18639v1 Announce Type: new Abstract: Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.
arXiv:2608.21606v1 Announce Type: new Abstract: Machine unlearning aims to remove the influence of targeted training data from a model while preserving its remaining capabilities, but evaluating whet...
arXiv:2608. 05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning.
The paper introduces FAB, an attack that uses meta‑learning to embed dormant adversarial behaviors into large language models (LLMs). These behaviors remain inactive until the model is finetuned by downstream users, at which point the model can exhibit unwanted actions such as unsolicited advertising, jailbreakability, or over‑refusal. FAB is shown to be effective across multiple LLMs and resilient to various finetuning settings.
The paper introduces FDCU, a dual‑constrained subspace projection framework designed to improve machine unlearning for large language models. FDCU limits parameter updates with a dual‑masking rule that preserves general knowledge via Fisher Information while preventing the activation of spurious suppressors through the Principle of Minimal Functional Intervention. Experiments show that FDCU achieves state‑of‑the‑art robustness against retraining attacks while maintaining near‑lossless general utility, thereby ensuring durable safety for LLMs.