arXiv AI By Dmitry Manning-Coe, Thomas Read, Anna Soligo, Oliver Clive-Griffin, Chun-Hei Yip, Rajashree Agrawal, Jason Gross

Interactions Between Crosscoder Features: A Compact Proofs Perspective

Read the original on arXiv AI →

arXiv:2606. 09940v1 Announce Type: cross Abstract: Dictionary learning methods like Sparse Autoencoders (SAEs) and crosscoders attempt to explain a model by decomposing its activations into independent features.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.