arXiv Machine Learning By Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei

HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models

Read the original on arXiv Machine Learning →

arXiv:2608. 12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv Machine Learning
4d ago

How to Tame a Multi-Headed Hydra? Adaptive Multi-Category Safety Steering for Large Language Models

The paper introduces CAM-Steer, a Category‑Adaptive Multi‑category Safety Steering framework that estimates risk for each harm category by comparing hidden states to safe and unsafe prototypes. It then combines safety directions into a single steering vector and applies a rotation whose angle is set by the estimated risks, preserving the hidden‑state norm. Experiments on three LLM backbones and seven harm categories show that CAM‑Steer outperforms baselines in defense success rate, even when multiple harm categories co‑occur, with negligible inference overhead.

By Chenxi Wang, Ruiyang Huang, Li Huang, Yifan Wu