arXiv Machine Learning

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

arXiv:2608. 12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.

arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv Machine Learning
4d ago

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.

By Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu
arXiv Machine Learning
Jul 30

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

arXiv:2607. 27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand.

By Yongjian Guo, Wanlun Ma, Lingyu Shen, Xi Xiao, Sheng Wen
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

arXiv Computation and Language
Aug 25

Safety-Aligned Weights Are Not Enough: Refusal-Teacher-Guided Finetuning Enhances Safety and Downstream Performance under Harmful Finetuning Attacks

arXiv:2506.07356v3 Announce Type: replace Abstract: While Finetuning-as-a-Service (FaaS) enables customization of Large Language Models (LLMs) using user data, this service is vulnerable to safety de...

By Seokil Ham, Yubin Choi, Yujin Yang, Seungju Cho, Younghun Kim, Changick Kim
arXiv AI
Sep 2

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

The paper investigates why large language models lose safety alignment after benign fine‑tuning. It argues that safety alignment relies on a low‑rank, output‑routing geometry that becomes flatter during fine‑tuning, and that after only 100 benign examples this routing is sharpened in output‑side MLPs, leading to fragile safety while general performance remains relatively intact. Techniques like LoRA and ASAM can delay this collapse by reducing output‑side sharpness, but their effectiveness diminishes with larger fine‑tuning scales.

By Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang