arXiv AI By Daming Luo

When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control

Read the original on arXiv AI →

arXiv:2607. 10226v1 Announce Type: new Abstract: We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.