When Does Activation Steering Change What a Model Computes From?
Read the original on arXiv Machine Learning →The paper investigates whether modifying a model’s internal activations—through activation steering—actually changes the computational state used in subsequent processing or merely biases the computation toward a desired output. In a controlled state‑tracking task, editing a trace‑supervised register causes the model to apply the next operation to the edited state, confirming that the edit changes the state. However, in two large‑language‑model settings (Qwen and Llama), mean activation steering does not replicate the natural internal configuration used during task execution; steering vectors are far larger than typical natural changes and achieve only a fraction of the effect of full‑layer patching. Thus, a successful steering intervention does not necessarily reproduce the natural target activation at the intervention layer, and should be interpreted as a state change only when later computation actually uses the edited value in the intended semantic way.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.