arXiv Machine Learning

When Does Activation Steering Change What a Model Computes From?

The paper investigates whether modifying a model’s internal activations—through activation steering—actually changes the computational state used in subsequent processing or merely biases the computation toward a desired output. In a controlled state‑tracking task, editing a trace‑supervised register causes the model to apply the next operation to the edited state, confirming that the edit changes the state. However, in two large‑language‑model settings (Qwen and Llama), mean activation steering does not replicate the natural internal configuration used during task execution; steering vectors are far larger than typical natural changes and achieve only a fraction of the effect of full‑layer patching. Thus, a successful steering intervention does not necessarily reproduce the natural target activation at the intervention layer, and should be interpreted as a state change only when later computation actually uses the edited value in the intended semantic way.

arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
arXiv AI
6d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza
arXiv AI
Sep 4

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.

By Vijay Erramilli