What Does Activation Steering Control? Attribution Across Answer Encodings and Output-Sensitive Subspaces
Read the original on arXiv Computation and Language →The paper investigates what aspects of language model behavior are controlled by activation steering. By introducing Cross‑Encoding Steering Evaluation, the authors show that steering effects often follow the extraction index of answer identifiers rather than the semantic content of the answers, especially at deeper layers. They also find that a low‑rank output‑sensitive component captures most of this effect, and that different datasets (NormBank, MNLI, SC101) exhibit varying preferences for extraction‑index versus semantic‑label following.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.