arXiv:2606.04160v2 Announce Type: replace-cross
Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse harmful or disallowed requ...
By Anna C. Marbut, Daniel R. Olson, Travis J. Wheeler
arXiv:2603.13359v2 Announce Type: replace
Abstract: Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusa...
By Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda
arXiv:2608.30197v1 Announce Type: new
Abstract: Safety alignment is essential for deploying large language models, requiring systems to prevent harmful compliance while preserving helpfulness on beni...
By Hoejoon Kwon, Byeonggeuk Lim, Kahyeon Kim, YoungBin Kim
The paper addresses the problem of over‑refusal in safety‑aligned large language models, where benign instructions are incorrectly rejected. It identifies that a small set of hypersensitive safety heads in transformer attention misfire on hard‑safe prompts, causing abnormal attention entanglement and high‑entropy routing conflicts that block necessary attention to target entities. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training‑free inference framework that dynamically suppresses these hypersensitive heads and fuses logits from dual branches to restore trustworthy reasoning while preserving intrinsic safety performance.
By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
arXiv:2609.25049v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions....
By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
arXiv:2609.00760v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate saf...
By Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim