arXiv AI

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

arXiv:2608. 02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward white-box attacks that directly identify and disrupt internal safety neurons or routes.

arXiv Machine Learning
Jul 30

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

arXiv:2607. 27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand.

By Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen
arXiv AI
Sep 3

SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment

The paper introduces SEAL, a training-time, parameter‑efficient defense that attaches a plug‑and‑play adapter to the shared expert component of Mixture‑of‑Experts models, and SEAL++, which adds an orthogonal constraint to preserve existing safety subspaces. By leveraging the always‑activated shared expert, SEAL mitigates the structural vulnerability of sparse routing to adversarial manipulation, reducing attack success rates by up to 60% with minimal impact on model capability. The approach is evaluated across six attack scenarios involving harmful prompting, jailbreaks, malicious fine‑tuning, and neuron pruning.

By Qingyu Meng, Yiwei Zha, Jiahuan Pei, Koen Hindriks, Herbert Bos, Min Chen
arXiv AI
Aug 26

NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution

NeuronGuard is a fine‑tuning defense for large language models that hardens them against both jailbreak and neuron‑level attacks. It redistributes safety signals across many neurons by identifying safety‑critical ones with per‑layer linear classifiers, enforcing refusal behavior when those neurons are ablated, and applying KL‑divergence regularization for consistency. A randomized gradient projection preserves task performance, and the authors provide a formal guarantee that NeuronGuard lowers the attack success rate upper bound, with experiments showing near‑zero success rates across multiple models and attack strategies.

By Anjun Gao, Yueyang Quan, Yufei Xia, Zhuqing Liu, Minghong Fang
arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv AI
Sep 2

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

The paper investigates why large language models lose safety alignment after benign fine‑tuning. It argues that safety alignment relies on a low‑rank, output‑routing geometry that becomes flatter during fine‑tuning, and that after only 100 benign examples this routing is sharpened in output‑side MLPs, leading to fragile safety while general performance remains relatively intact. Techniques like LoRA and ASAM can delay this collapse by reducing output‑side sharpness, but their effectiveness diminishes with larger fine‑tuning scales.

By Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang