OpenAI Blog

Toward understanding and preventing misalignment generalization

Read the original on OpenAI Blog →

We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.

Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.

arXiv Machine Learning
Jun 5

In-Training Defenses against Emergent Misalignment in Language Models

arXiv:2508. 06249v3 Announce Type: replace Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain.

By David Kacz\'er, Magnus J{\o}rgenv{\aa}g, Clemens Vetter, Esha Afzal, Robin Haselhorst, Lucie Flek, Florian Mai