arXiv AI By Evan Duan

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects

Read the original on arXiv AI →

arXiv:2606. 08365v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are increasingly used to steer language models, but feature steering is rarely clean: the same intervention can behave inconsistently across contexts and perturb unrelated features.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 4

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.

By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
arXiv Machine Learning
Jul 28

Steering grids for sparse-autoencoder features: when a top-context label names an activation regime rather than a causal axis

arXiv:2605. 03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude.

By Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen