arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.
By Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
arXiv:2606. 31591v1 Announce Type: cross Abstract: Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insecure code, leads to broadly misaligned behaviour on unrelated prompts.
By Jason R. Brown, Patrick Leask, Lev McKinney
arXiv:2607. 26654v1 Announce Type: cross Abstract: Post-training alignment is often shallow, eroding under fine-tuning.
By Desiree Cho, Cameron Tice, Bernie Hogan, Hunar Batra, Puria Radmard, Jun Zhao, Nigel Shadbolt
arXiv:2607. 25063v1 Announce Type: new Abstract: Developers judge a model checkpoint by how it behaves.
By Cen Lu, Yung-Chen Tang, Andrea Cavallaro
arXiv:2608. 03887v1 Announce Type: new Abstract: Fine-tuning a large language model on new data degrades what it previously learned.
By Alberto Acedo
arXiv:2606. 07596v1 Announce Type: new Abstract: Fine-tuning often introduces spurious correlations alongside task knowledge, causing systematic failures on underrepresented groups.
By Edward Sun, Dmitrii Troitskii
arXiv:2608. 17162v1 Announce Type: new Abstract: What a language model internalizes from fine-tuning is usually diagnosed after the fact.
By Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang
arXiv:2607. 26389v1 Announce Type: cross Abstract: Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated.
By Hasibur Rahman, Smit Desai
arXiv:2608. 08829v1 Announce Type: cross Abstract: Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice fixes the injection layers globally per task.
By Muhammad Faishal Adly Nelwan, Alfan Farizki Wicaksono
arXiv:2608. 04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence.
By Kotaro Yoshida, Laura Gomezjurado Gonzalez, Yukinori Yamamoto, Yuji Naraki, Ryotaro Shimizu, Wenya Wang
arXiv:2607. 18293v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) teaches large language models new skills through a teacher that shares the student's backbone and supervises its own rollouts.
By Yingzi Ma, Zichen Zhu, Ming Jiang, Chaowei Xiao
arXiv:2601. 18699v2 Announce Type: replace Abstract: Sequential fine-tuning of Large Language Models (LLMs) adaptation to target tasks often triggers catastrophic forgetting, where the acquisition of novel target skills degrades ancestral capabilities.
By Gustav Olaf Yunus Laitinen-Fredriksson Lundstrom-Imanov