K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models
Read the original on arXiv Computation and Language →The paper investigates K/V-cache interventions—transplanting a target-conditioned key/value trajectory into a source-persona generation—as a method for controlling persona in decoder-only language models. Experiments on Llama‑3.1‑8B across 13 configurations reveal that strong representation alignment (measured by V‑gap) does not guarantee behavioral persona expression, with only mid‑layer replacements achieving both alignment and lexical diversity. Position perturbations uniformly suppress persona expression, highlighting that representation similarity alone is insufficient to predict downstream behavior.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.