arXiv Computation and Language By Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han

K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models

Read the original on arXiv Computation and Language →

The paper investigates K/V-cache interventions—transplanting a target-conditioned key/value trajectory into a source-persona generation—as a method for controlling persona in decoder-only language models. Experiments on Llama‑3.1‑8B across 13 configurations reveal that strong representation alignment (measured by V‑gap) does not guarantee behavioral persona expression, with only mid‑layer replacements achieving both alignment and lexical diversity. Position perturbations uniformly suppress persona expression, highlighting that representation similarity alone is insufficient to predict downstream behavior.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 14

Synthetic Persona Pretraining: Alignment from Token Zero

arXiv:2608. 13482v1 Announce Type: cross Abstract: As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical.

By Julian Minder, Viktor Moskvoretskii, Raghav Singhal, Difan Jiao, Andy Arditi, Shaobo Cui, Yiderigun Borjigin, Kartik Bali, Stefan Krsteski, Harsh Raj, Huu Nguyen, Jannik Brinkmann, Ashton Anderson, Roland Aydin, Robert West