arXiv AI By Peiyang Liu, Xi Wang, Ziqiang Cui, Di Liang, Wei Ye

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

Read the original on arXiv AI →

arXiv:2608. 08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.