arXiv AI By Qinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang, Dongxia Wang

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

Read the original on arXiv AI →

arXiv:2606. 03601v1 Announce Type: cross Abstract: While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.