arXiv AI

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration

arXiv Computation and Language
Sep 23

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

The paper addresses the problem of over‑refusal in safety‑aligned large language models, where benign instructions are incorrectly rejected. It identifies that a small set of hypersensitive safety heads in transformer attention misfire on hard‑safe prompts, causing abnormal attention entanglement and high‑entropy routing conflicts that block necessary attention to target entities. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training‑free inference framework that dynamically suppresses these hypersensitive heads and fuses logits from dual branches to restore trustworthy reasoning while preserving intrinsic safety performance.

By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.

arXiv AI
Sep 2

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

The paper investigates why large language models lose safety alignment after benign fine‑tuning. It argues that safety alignment relies on a low‑rank, output‑routing geometry that becomes flatter during fine‑tuning, and that after only 100 benign examples this routing is sharpened in output‑side MLPs, leading to fragile safety while general performance remains relatively intact. Techniques like LoRA and ASAM can delay this collapse by reducing output‑side sharpness, but their effectiveness diminishes with larger fine‑tuning scales.

By Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang