arXiv:2508. 16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts.
By David Chanin, Adri\`a Garriga-Alonso
arXiv:2602. 14687v2 Announce Type: replace-cross Abstract: Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations.
By David Chanin, Adri\`a Garriga-Alonso
The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.
By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv:2609.06557v1 Announce Type: new
Abstract: Large language models (LLMs) are often considered fragile under aggressive sparsification, and maintaining reliable performance typically requires stic...
By Hyeondo Jang, Kwanhee Lee, Dongyeop Lee, Namhoon Lee
arXiv:2609.15064v1 Announce Type: new
Abstract: Reinforcement learning (RL) is widely utilized in large language model training to improve targeted capabilities, yet how RL reshapes a model remains p...
By Lingheng Du, Yiming Tang, Xufeng Duan, Dianbo Liu
arXiv:2606. 26620v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features.
By XinYang He, Wei Wang, Bing Zhao, Xuan Ren, WenBo Li, WeiXu Qiao, Hu Wei, Lin Qu