arXiv AI

Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment

arXiv:2606. 23700v1 Announce Type: cross Abstract: Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM operates through disruption of the model's aligned character rather than direct learning of harmful content.

arXiv AI
Sep 16

TAME: Token Attribution and Masking for Emergent misalignment

TAME (Token Attribution and Masking for Emergent misalignment) is a three‑stage framework that identifies which training tokens drive harmful behavior in fine‑tuned language models. It first scores tokens by how much fine‑tuning increases their likelihood, then characterizes patterns among high‑attribution tokens, and finally validates them by masking during training. Experiments on Llama and Qwen show that masking the top‑attribution tokens reduces emergent misalignment by 23‑ to 36‑fold, while random masking has no effect.

By Md Rayhanul Masud, Md Rizwan Parvez
arXiv Computation and Language
Sep 11

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models

The paper investigates K/V-cache interventions—transplanting a target-conditioned key/value trajectory into a source-persona generation—as a method for controlling persona in decoder-only language models. Experiments on Llama‑3.1‑8B across 13 configurations reveal that strong representation alignment (measured by V‑gap) does not guarantee behavioral persona expression, with only mid‑layer replacements achieving both alignment and lexical diversity. Position perturbations uniformly suppress persona expression, highlighting that representation similarity alone is insufficient to predict downstream behavior.

By Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han