arXiv AI By Viola Zhong, Qirui Li

Refusal Lives Downstream of Persona in Chat Models

Read the original on arXiv AI →

arXiv:2606. 26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.