arXiv Machine Learning

An Empirical Study of Feature Selection Granularity

arXiv:2607. 24145v1 Announce Type: new Abstract: Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task.

arXiv Machine Learning
Aug 4

Beyond Noise: A Hypothesis Testing Approach to Robust Feature Selection

arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.

By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv Machine Learning
Aug 13

Towards Truly Unsupervised Evaluation of Feature Selection

arXiv:2608. 12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method.

By Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek
arXiv Machine Learning
Jul 21

MinShap: A Shapley-Based Framework for Feature Redundancy

arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.

By Chenghui Zheng, Garvesh Raskutti
arXiv Machine Learning
Sep 18

Null importance: Disentangling relevance for interpretable machine learning

The paper introduces a unified framework called null importance to clarify different notions of feature relevance in interpretable machine learning. It defines null importance at the population level for various relevance concepts—marginal, conditional, predictive risk, functional invariance, and causal effects—and demonstrates how each answers distinct scientific questions. Through theoretical analysis, simulations, and case studies on fairness and genomic modeling, the authors show when these null notions coincide or diverge and how different importance methods target them.

By Garvesh Raskutti, Kris Sankaran, Jiaxin Ye