arXiv Machine Learning By Yuki Nakamura

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol

Read the original on arXiv Machine Learning →

arXiv:2605. 24583v3 Announce Type: replace Abstract: Comparing a model's internal activations before and after alignment is a natural way to ask what safety training changes: one forms the matrix of paired aligned-minus-base activations on safety-relevant inputs and reads off its effective rank or top direction.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.