arXiv:2603.13359v2 Announce Type: replace
Abstract: Language models are commonly fine-tuned for safety alignment to refuse harmful prompts. One approach fine-tunes them to generate categorical refusa...
By Rishab Alagharu, Ishneet Sukhvinder Singh, Shaibi Shamsudeen, Zhen Wu, Ashwinee Panda
arXiv:2609.25602v1 Announce Type: new
Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
By Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh
arXiv:2609.14759v1 Announce Type: cross
Abstract: Alignment applied after pretraining is shallow in a measurable way: a single direction in a model's residual stream can be edited out, and the model...
By Orion Reblitz-Richardson
arXiv:2603.27518v4 Announce Type: replace
Abstract: Aligned language models that are trained to refuse harmful requests also exhibit over-refusal: they decline safe instructions that seemingly resemb...
By Utsav Maskey, Mark Dras, Usman Naseem
arXiv:2606. 13720v1 Announce Type: new Abstract: Arditi et al.
By Elisabetta Rocchetti, Alfio Ferrara
arXiv:2608.29070v1 Announce Type: new
Abstract: Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in pri...
By Zimo Shi, Xander Tifft, Wen Xing