arXiv Machine Learning

Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol

arXiv:2605. 24583v3 Announce Type: replace Abstract: Comparing a model's internal activations before and after alignment is a natural way to ask what safety training changes: one forms the matrix of paired aligned-minus-base activations on safety-relevant inputs and reads off its effective rank or top direction.

arXiv Computation and Language
Sep 11

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models

The paper investigates K/V-cache interventions—transplanting a target-conditioned key/value trajectory into a source-persona generation—as a method for controlling persona in decoder-only language models. Experiments on Llama‑3.1‑8B across 13 configurations reveal that strong representation alignment (measured by V‑gap) does not guarantee behavioral persona expression, with only mid‑layer replacements achieving both alignment and lexical diversity. Position perturbations uniformly suppress persona expression, highlighting that representation similarity alone is insufficient to predict downstream behavior.

By Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han
arXiv AI
Jul 21

How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?

arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.

By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin