arXiv Machine Learning By Dibyanayan Bandyopadhyay, Asif Ekbal

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

Read the original on arXiv Machine Learning →

arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.