arXiv:2310.16295v2 Announce Type: replace-cross
Abstract: Neural network have achieved remarkable successes in many scientific fields. However, the interpretability of the neural network model is sti...
By Zhimin Li, Shusen Liu, Kailkhura Bhavya, Peer-Timo Bremer, Valerio Pascucci
arXiv:2607. 08641v1 Announce Type: new Abstract: Over the last few years, there has been an increased interest in making machine learning models more interpretable.
By Yann Claes, Pierre Geurts, V\^an Anh Huynh-Thu
The paper introduces an importance‑scoring metric for multi‑head transformer attention heads applied to tabular data, a domain where transformers have been less studied. Experiments on 40 diverse tabular datasets show that removing heads with the lowest importance scores has minimal impact on performance, while removing the most important head first causes the largest drop. The study finds that important heads are distributed across layers and vary significantly across different tabular schemas, suggesting that the proposed score can help reduce redundancy and improve transformer efficiency.
By Ahmad Jad Allah, Kazi F. Akhter, Md. Kamrozzaman Bhuiyan, Manar D. Samad
arXiv:2609.24126v1 Announce Type: cross
Abstract: Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important fe...
By Xuhui Liu, Lili Zheng
arXiv:2606. 16939v1 Announce Type: cross Abstract: A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly produce model behavior.
By Naiyu Yin, Dennis Wei, Tian Gao, Amit Dhurandhar, Karthikeyan Natesan Ramamurthy, Yue Yu
arXiv:2505.12683v2 Announce Type: replace
Abstract: Key feature fields need bigger embedding dimensionality, others need smaller. This demands automated dimension allocation. Existing approaches, suc...
By Yihong Huang, Chen Chu
arXiv:2511.15371v3 Announce Type: replace
Abstract: Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous m...
By Eddie Conti, \'Alvaro Parafita, Axel Brando
arXiv:2606. 29951v1 Announce Type: new Abstract: Interpretable Mesomorphic Neural Networks (IMNs) offer a promising framework that combines the predictive power of deep neural networks with the interpretability of linear models.
By Hugo L. Hammer, Vajira Thambawita, Kristoffer Herland Hellton, P{\aa}l Halvorsen
arXiv:2607. 25529v1 Announce Type: new Abstract: As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability.
By Qitao Chen, Dongfu Yin, F. Richard Yu
The paper introduces a flexible symbolic framework that efficiently computes logical explanations for deep neural networks by parameterizing explanations with internal neuron activations and leveraging general-purpose logical engines like SMT solvers. Unlike previous methods that rely on specialized verifiers or are limited to individual input features, this approach is not restricted in shape and can scale to deep architectures. Experiments on image recognition and medical benchmarks demonstrate improved computational efficiency and the ability to explain networks that were previously intractable for logic-based methods.
By Tom\'a\v{s} Kol\'arik, Faezeh Labbaf, Fabrizio Leopardi, Grigory Fedyukovich, Michael Wand, Natasha Sharygina
The paper introduces eXplaining to Learn (eX2L), an interpretable framework that regularizes a classifier by penalizing similarity between Grad‑CAM maps of the main label classifier and a confounder classifier. This approach decorrelates confounding features from latent representations during training. On the Spawrious Many‑to‑Many Hard Challenge benchmark, eX2L outperforms the current state‑of‑the‑art by 5.49% in average accuracy and 10.90% in worst‑group accuracy, while also demonstrating functional domain invariance through explicit label‑nuisance decoupling.
By Paulo Mario P. Medina, Jose Marie Antonio Mi\~noza, Sebastian C. Iba\~nez
arXiv:2607. 12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models.
By Ayush Karmacharya (Purdue University), Luke Luschwitz (Purdue University), Lucia Romero (Purdue University), Yanan Niu (EPFL), Joseph Campbell (Purdue University)