arXiv AI

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

arXiv:2606. 18322v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features.