arXiv AI By Ash Manvi, Samreena Tajreen

Finding Usable Weight Mechanisms with Tiled SVD

Read the original on arXiv AI →

arXiv:2608. 06969v1 Announce Type: new Abstract: The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-activating text.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.