Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint.
arXiv:2607. 11475v1 Announce Type: new Abstract: Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance.
arXiv:2609.36569v1 Announce Type: cross Abstract: Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fix...
The paper investigates how fine‑tuning large language models with a small number of harmful examples can erode their refusal behavior, and explores whether localizing safety‑related behavior to specific layers or directions can provide robust defenses. Experiments across six checkpoints from four model families show that harmful and benign prompts remain linearly separable after attack, and that patching clean hidden states or freezing layers up to a transition depth can restore refusal. However, attackers can bypass these defenses by spreading updates or targeting singular directions, indicating that adaptive fine‑tuning can defeat localized repairs and highlighting the need for multiple defensive checks.
arXiv:2609.13714v1 Announce Type: new Abstract: An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to no...
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.