arXiv Machine Learning By Muhammad Khan, Shlok Channawar, Akshaj Gurugubelli, Girish Gupta, Aditya Shah

Look Before You Steer: Geometry Predicts SAE Feature Steerability

Read the original on arXiv Machine Learning →

The paper investigates whether properties of a Self‑Attention Encoder (SAE) that can be computed before any forward pass can predict how difficult it is to steer individual SAE features. It finds that decoder‑space geometry—specifically neighbor density and maximum cosine similarity to nearby decoder directions—partially predicts feature steerability, with significant correlations across multiple model scales, widths, and architectures. The relationship holds up to a certain layer depth, suggesting a boundary where steering becomes too costly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 23

Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models

arXiv:2607. 19364v1 Announce Type: new Abstract: Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering methods often rely on learned steering objectives or single-criterion feature selection.

By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan
arXiv Machine Learning
Jun 11

When is Your LLM Steerable?

arXiv:2606. 11599v1 Announce Type: cross Abstract: Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration.

By Chenrui Fan, Yize Cheng, Ming Li, Soheil Feizi, Tianyi Zhou
Hugging Face Trending Papers
Jun 10

When is Your LLM Steerable?

Activation steering offers a lightweight approach to control language models' behavior at inference time, but whether it succeeds or fails heavily depends on the prompt, concept, model, and steering configuration. Finding the regime and boundaries of successful steering typically requires expensive grid searches and post-hoc evaluation of full autoregressive rollouts.

arXiv Machine Learning
Sep 1

World Model Control by Trajectory Reachability Metrics

arXiv:2605.22164v2 Announce Type: replace Abstract: Latent world models can learn representations that contain information needed for control, while the downstream controller may still rank candidate...

By Liangyu Li, Shengzhi Wang, Libin Qiu, Mingliang Xiong, Qingwen Liu