arXiv AI

Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updates

Hugging Face Trending Papers
Jul 13

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint.

arXiv Machine Learning
4d ago

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

The paper investigates how fine‑tuning large language models with a small number of harmful examples can erode their refusal behavior, and explores whether localizing safety‑related behavior to specific layers or directions can provide robust defenses. Experiments across six checkpoints from four model families show that harmful and benign prompts remain linearly separable after attack, and that patching clean hidden states or freezing layers up to a transition depth can restore refusal. However, attackers can bypass these defenses by spreading updates or targeting singular directions, indicating that adaptive fine‑tuning can defeat localized repairs and highlighting the need for multiple defensive checks.

By Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim
arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov