arXiv Machine Learning By Benjamin Shih, John Winnicki, Eric Darve

When Does Activation Steering Change What a Model Computes From?

Read the original on arXiv Machine Learning →

The paper investigates whether modifying a model’s internal activations—through activation steering—actually changes the computational state used in subsequent processing or merely biases the computation toward a desired output. In a controlled state‑tracking task, editing a trace‑supervised register causes the model to apply the next operation to the edited state, confirming that the edit changes the state. However, in two large‑language‑model settings (Qwen and Llama), mean activation steering does not replicate the natural internal configuration used during task execution; steering vectors are far larger than typical natural changes and achieve only a fraction of the effect of full‑layer patching. Thus, a successful steering intervention does not necessarily reproduce the natural target activation at the intervention layer, and should be interpreted as a state change only when later computation actually uses the edited value in the intended semantic way.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron