arXiv Machine Learning

Emergent Misalignment Recruits a Pre-existing Persona Subspace

arXiv:2607. 21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment.