arXiv Computation and Language
Aug 25

What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces

The paper investigates what aspects of language model behavior are controlled by activation steering. By introducing Cross‑Encoding Steering Evaluation, the authors show that steering effects often follow the extraction index of answer identifiers rather than the semantic content of the answers, especially at deeper layers. They also find that a low‑rank output‑sensitive component captures most of this effect, and that different datasets (NormBank, MNLI, SC101) exhibit varying preferences for extraction‑index versus semantic‑label following.

By Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki