arXiv:2609.37857v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. Howev...
By Zhenting Huang, Junnan Liu, Qianren Mao, Zhixing Tan, Bo Jiang
arXiv:2606. 27321v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their polysemantic activations into a larger set of sparse, more monosemantic features.
By Nathana\"el Jacquier, Maria Vakalopoulou, Mahdi S. Hosseini
The paper introduces MASS, a hierarchical data selection method that first groups data using low‑dimensional manifold coordinates learned by a dense autoencoder, then applies a TopK sparse autoencoder for quality‑aware feature coverage within each group. This approach addresses the shortcomings of traditional diversity metrics that mix semantic, supervisory, and noise signals. Experiments on Vision Flan and LLaVA‑CoT demonstrate that MASS outperforms existing baselines across various budgets and can match or exceed full‑data training with only a small subset.
By Peng Sun, Yi Yang, Antong Zhang, Chunxiao Li, Yanbo Wang, Dianbo Liu, xin chen, Kai Yu, Lu Chen, Tianfan Fu
ShiftSplit-AD is a method that separates domain shift from defects in visual anomaly detection by decomposing the residual matrix of DINOv2 features into low‑rank and row‑sparse components. The sparse component is used for scoring anomalies, optionally fused with the low‑rank part. Experiments on AeBAD‑S show that sparse‑only scoring raises image AUROC from 0.6780 to 0.7294 and AUPRC from 0.8052 to 0.8465, but it also lowers clean AUROC on MVTec categories and hurts Bottle localization, highlighting a trade‑off between filtering shift and preserving defect information.
By Muhamathu Ameer Ali Aacaas Muhamath
arXiv:2508. 16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts.
By David Chanin, Adri\`a Garriga-Alonso
arXiv:2608.28806v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) disentangle model activations into interpretable features and are widely used for steering large language models. Most exist...
By Yutian Liu, Xu Wang, Difan Zou