arXiv AI By Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

Read the original on arXiv AI →

arXiv:2608. 15772v1 Announce Type: new Abstract: When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.