arXiv AI

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

arXiv:2607. 10226v1 Announce Type: new Abstract: We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior.

arXiv Machine Learning
Jul 28

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

arXiv:2605. 03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude.

By Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen