arXiv AI

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.

arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
arXiv AI
Sep 10

DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment

The paper introduces DSPA, a dynamic sparse autoencoder (SAE) steering technique that aligns language model outputs with user preferences during inference, avoiding costly weight updates. DSPA constructs a conditional-difference map from preference triples to adjust token-active latents, improving MT‑Bench scores and matching AlpacaEval performance on models like Gemma‑2 and Qwen3 while preserving accuracy. It demonstrates robustness with limited preference data, outperforms the two‑stage RAHF‑SCIT pipeline in FLOPs, and reveals that preference directions are largely driven by discourse and stylistic cues.

By James Wedgwood, Aashiq Muhamed, Mona T. Diab, Virginia Smith
arXiv Machine Learning
Sep 22

Look Before You Steer: Geometry Predicts SAE Feature Steerability

The paper investigates whether properties of a Self‑Attention Encoder (SAE) that can be computed before any forward pass can predict how difficult it is to steer individual SAE features. It finds that decoder‑space geometry—specifically neighbor density and maximum cosine similarity to nearby decoder directions—partially predicts feature steerability, with significant correlations across multiple model scales, widths, and architectures. The relationship holds up to a certain layer depth, suggesting a boundary where steering becomes too costly.

By Muhammad Khan, Shlok Channawar, Akshaj Gurugubelli, Girish Gupta, Aditya Shah
arXiv AI
Sep 10

Steering Geometry: Validating Human Value Geometry in LLM Steering Space

arXiv:2609.06289v1 Announce Type: cross Abstract: As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inferenc...

By Mohammad Mahdi Abootorabi, Armin Saghafian, Ali Bazshoushtari, Hamid Rezaei, EunJeong Hwang, Vered Shwartz, Parvin Mousavi, Purang Abolmaesumi
Hugging Face Trending Papers
Jun 10

When is Your LLM Steerable?

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.

arXiv Machine Learning
Jun 15

Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions

arXiv:2605. 05983v2 Announce Type: replace Abstract: Recently, steering vectors (SVs) have emerged as an effective and lightweight approach to steer behaviors of large language models (LLMs), among which fine-tuned SVs are more effective than optimization-free ones.

By Yuntai Bao, Qinfeng Li, Xinyan Yu, Ge Su, Wenqi Zhang, Liu Yan, Haiqin Weng, Jianwei Yin, Xuhong Zhang