arXiv AI

SAILS: Surrogate-based Analysis of Interactions via Local Effect Smooths

arXiv:2606. 09404v1 Announce Type: cross Abstract: Feature interactions drive much of the predictive power of machine learning models, yet existing explanation methods only detect and quantify interactions without revealing their functional form, or visualize only restricted interaction types.

arXiv Machine Learning
Jun 10

Interpretable deep convolutional model for nonlinear multivariate time series in complex systems

arXiv:2501. 04339v2 Announce Type: replace-cross Abstract: We introduce the Deep Convolutional Interpreter for Time Series (DCIts), a deep-learning architecture for nonlinear multivariate time series that provides sample-specific, locally interpretable descriptions of the underlying interaction structure.

By Domjan Baric, Davor Horvatic
Hugging Face Trending Papers
Jul 27

Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended direction; and activation-based feature selection can miss features that produce the desired output change.