arXiv Machine Learning By Mohammed Suhail B Nadaf

Emergent Misalignment Recruits a Pre-existing Persona Subspace

Read the original on arXiv Machine Learning →

arXiv:2607. 21356v1 Announce Type: new Abstract: Fine-tuning an aligned language model on a narrow stream of bad advice can make it broadly misaligned on questions unrelated to the training data, a phenomenon called emergent misalignment.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.