arXiv Machine Learning

Look Before You Steer: Geometry Predicts SAE Feature Steerability

The paper investigates whether properties of a Self‑Attention Encoder (SAE) that can be computed before any forward pass can predict how difficult it is to steer individual SAE features. It finds that decoder‑space geometry—specifically neighbor density and maximum cosine similarity to nearby decoder directions—partially predicts feature steerability, with significant correlations across multiple model scales, widths, and architectures. The relationship holds up to a certain layer depth, suggesting a boundary where steering becomes too costly.

arXiv AI
Jul 23

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.

By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
Hugging Face Trending Papers
Jun 10

When is Your LLM Steerable?

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.

arXiv Machine Learning
Sep 1

World Model Control by Trajectory Reachability Metrics

arXiv:2605.22164v2 Announce Type: replace Abstract: Latent world models can learn representations that contain information needed for control, while the downstream controller may still rank candidate...

By Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu
arXiv AI
Sep 10

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

arXiv:2609.06289v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...

By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
arXiv Machine Learning
Jul 28

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

arXiv:2605. 03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude.

By Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen
arXiv Machine Learning
Aug 27

Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.

By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron