arXiv AI By Yu Feng, Chunting Zang, Chen Shen, Rui Miao, Ge Teng, Weidong Cai, Jieping Ye

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

Read the original on arXiv AI →

arXiv:2608. 03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.