arXiv Machine Learning

AsFT: Anchoring Safety During LLM Fine-Tuning Within Narrow Safety Basin

arXiv:2506. 08473v4 Announce Type: replace Abstract: Fine-tuning large language models (LLMs) improves performance but introduces critical safety vulnerabilities: even minimal harmful data can severely compromise safety measures.

arXiv AI
Aug 24

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

CLEAR is a conditional safety adaptation framework that employs a lightweight hidden‑state gate to continuously control the activation strength of a safety low‑rank adapter. It aims to reduce harmful completions while preserving the performance of the frozen backbone on benign prompts. Experiments on safety and utility benchmarks, including HarmBench and GSM8K, show that CLEAR significantly lowers HarmBench ASR and improves utility compared to globally applied safety tuning methods such as SFT or standard LoRA.

By Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo