arXiv AI By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 23

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

The paper addresses the problem of over‑refusal in safety‑aligned large language models, where benign instructions are incorrectly rejected. It identifies that a small set of hypersensitive safety heads in transformer attention misfire on hard‑safe prompts, causing abnormal attention entanglement and high‑entropy routing conflicts that block necessary attention to target entities. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training‑free inference framework that dynamically suppresses these hypersensitive heads and fuses logits from dual branches to restore trustworthy reasoning while preserving intrinsic safety performance.

By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
arXiv Computation and Language
Aug 25

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.

By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.