arXiv Machine Learning

Equivariant Sparse Autoencoders: Mechanistic Interpretability of Neural Networks on Symmetric Data

arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.

Hugging Face Trending Papers
Aug 12

Reducing Symmetry Increase in Equivariant Neural Networks

Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.

arXiv Machine Learning
Jun 18

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.

By Dibyanayan Bandyopadhyay, Asif Ekbal
arXiv Machine Learning
Jun 24

Similarity of Neural Network Representations in Superposition

arXiv:2604. 00208v2 Announce Type: replace Abstract: Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features.

By Sunny Liu, Habon Issa, Andr\'e Longon, Liv Gorton, Meenakshi Khosla, Alex Williams, David Klindt