arXiv Machine Learning By Tal Ellinson, Hadi Mohasel Afshar, Sally Cripps

Hide&Seek: Learning to Explain in an End-to-End Differentiable Network

Read the original on arXiv Machine Learning →

arXiv:2608. 16689v1 Announce Type: cross Abstract: Instance-wise feature selection is a valuable tool for interpreting labeled data and the predictions of black-box models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
Aug 25

Interpretable AI with Local Distillation

Interpretable AI with Local Distillation proposes a method where a black‑box teacher model guides a regularized linear student model at each query point. The teacher defines locality by upweighting training observations with similar predicted outcomes and anchors the fit with its own prediction at the query point, treated as a pseudo‑observation. By adding Gaussian randomization and refitting, the approach identifies reliable features and stable subgroups, achieving near‑teacher accuracy while producing sparse, locally interpretable linear models.

By Erin Craig, Yiling Huang, Snigdha Panigrahi
arXiv AI
Sep 10

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.

By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith