Steering Interference Reflects the Model's Defaults, Not the Behavior Directions
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 19831v1 Announce Type: cross Abstract: Aligned language models gate behaviors such as refusal and language routing through sparse feed forward neurons, yet no theory predicts when a single neuron intervention controls a behavior coherently rather than collapsing the output.
arXiv:2608. 06578v1 Announce Type: new Abstract: Frontier language models are trained using distinct data, objectives, and safety pipelines.
arXiv:2606. 24952v1 Announce Type: cross Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should be able to modify it.
The paper investigates how steering a language model along a trait vector—specifically for refusal and sycophancy—affects its internal computation. By splitting the model’s computation into a reconstruction circuit before the trait vector and a transmission circuit after it, the authors find that for refusal both circuits are compact and faithful, with the transmission circuit alone largely restoring the lost refusal signal. For sycophancy, the transmission circuit remains compact but the reconstruction circuit is broader and only partially faithful, indicating that the circuits used for steering are not necessarily the same as those that generate the behavior.
arXiv:2609.25602v1 Announce Type: new Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
arXiv:2606.04160v2 Announce Type: replace-cross Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse harmful or disallowed requ...