Hugging Face Trending Papers

HyperSafe: Inference-Time Safety Recovery for Fine-Tuned Language Models

Read the original on Hugging Face Trending Papers →

Safety alignment in large language models can be fragile under fine-tuning, as even benign task adaptation may increase harmful compliance. Existing defenses mainly follow two directions: they either intervene during or after fine-tuning through retraining or weight modification, which can be costly and may hurt task performance, or they use model-agnostic safety classifiers, which may miss failures specific to a given fine-tuned checkpoint.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.