arXiv Machine Learning

Hide&Seek: Learning to Explain in an End-to-End Differentiable Network

arXiv:2608. 16689v1 Announce Type: cross Abstract: Instance-wise feature selection is a valuable tool for interpreting labeled data and the predictions of black-box models.

arXiv Statistics ML
Aug 25

Interpretable AI with Local Distillation

Interpretable AI with Local Distillation proposes a method where a black‑box teacher model guides a regularized linear student model at each query point. The teacher defines locality by upweighting training observations with similar predicted outcomes and anchors the fit with its own prediction at the query point, treated as a pseudo‑observation. By adding Gaussian randomization and refitting, the approach identifies reliable features and stable subgroups, achieving near‑teacher accuracy while producing sparse, locally interpretable linear models.

By Erin Craig, Yiling Huang, Snigdha Panigrahi
arXiv AI
Sep 10

SAEs Can Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs

The paper introduces Dynamic DAE Guardrails (DSG), a method that uses Dynamic Sparse Autoencoders to perform precision unlearning in large language models. DSG leverages principled feature selection and a dynamic classifier to target activation-based unlearning, outperforming existing gradient‑based methods in terms of computational efficiency, stability, sequential unlearning, resistance to relearning attacks, data efficiency, and interpretability.

By Aashiq Muhamed, Jacopo Bonato, Mona Diab, Virginia Smith
arXiv AI
Aug 19

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

The paper proposes a Learnware-based framework for deploying scene‑specific CSI feedback models in 6G systems. A centralized AI data center maintains a catalog of pre‑trained models, each tagged with semantic and statistical specifications. Base stations retrieve the most relevant model using only statistical fingerprints, which reduces data privacy risks, lowers retrieval latency, and cuts fine‑tuning effort, achieving up to 57.7% performance gains over a general model.

By Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng, Shi Jin, Zhi-Hua Zhou
arXiv AI
Aug 13

Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection

arXiv:2512. 13374v2 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether their internal representations capture problem structure or algorithmic behavior.

By Francesca Da Ros, Luca Di Gaspero, Kevin Roitero