arXiv Machine Learning By Jiaming Li, Haoran Ye, Yukun Chen, Xinyue Li, Lei Zhang, Hamid Alinejad-Rokny, Jimmy Chih-Hsien Peng, Min Yang

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models

Read the original on arXiv Machine Learning →

arXiv:2506. 07691v2 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) are a cornerstone of mechanistic interpretability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 10

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.

By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith