arXiv AI

When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models

arXiv:2608. 03201v1 Announce Type: new Abstract: Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs.