arXiv AI

Where Steering Signals Come From: Activation Source Selection in Activation Steering

arXiv:2607. 25270v1 Announce Type: cross Abstract: Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of these steering signals is often treated as a secondary detail.

arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
Hugging Face Trending Papers
Jun 10

When is Your LLM Steerable?

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.

arXiv AI
Sep 2

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

The paper investigates how steering vectors influence large language models (LLMs) by conducting a mechanistic case study on refusal behavior. Using a multi-token activation patching framework, the authors find that steering methods primarily target the OV circuit of the attention mechanism, largely ignoring the QK circuit, and that these circuits are functionally interchangeable across different steering approaches. The study also shows that steering vectors can be sparsified by 85–96% with minimal performance loss and that key dimensions are consistently identified across methods.

By Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha