arXiv AI

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

The paper investigates why large language models lose safety alignment after benign fine‑tuning. It argues that safety alignment relies on a low‑rank, output‑routing geometry that becomes flatter during fine‑tuning, and that after only 100 benign examples this routing is sharpened in output‑side MLPs, leading to fragile safety while general performance remains relatively intact. Techniques like LoRA and ASAM can delay this collapse by reducing output‑side sharpness, but their effectiveness diminishes with larger fine‑tuning scales.

arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv AI
Aug 24

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

CLEAR is a conditional safety adaptation framework that employs a lightweight hidden‑state gate to continuously control the activation strength of a safety low‑rank adapter. It aims to reduce harmful completions while preserving the performance of the frozen backbone on benign prompts. Experiments on safety and utility benchmarks, including HarmBench and GSM8K, show that CLEAR significantly lowers HarmBench ASR and improves utility compared to globally applied safety tuning methods such as SFT or standard LoRA.

By Chengxiao Wang, Enyi Jiang, Xiaojing Liao, Sanmi Koyejo