arXiv Machine Learning By Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan, Neel Nanda

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Read the original on arXiv Machine Learning →

arXiv:2606. 26071v1 Announce Type: new Abstract: A central goal of safety research is determining whether a model is misaligned.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.