arXiv:2607. 24145v1 Announce Type: new Abstract: Feature selection aims to identify the most informative and relevant features for a given dataset, either in terms of capturing the underlying data structure and distribution better, or with respect to the performance on a downstream task.
By Muhammad Rajabinasab, Arthur Zimek
arXiv:2609.24126v1 Announce Type: cross
Abstract: Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important fe...
By Xuhui Liu, Lili Zheng
arXiv:2512. 11081v2 Announce Type: replace-cross Abstract: Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models.
By Kata Vuk, Nicolas Alexander Ihlo, Merle Behr
arXiv:2508. 14268v2 Announce Type: replace-cross Abstract: Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2608. 04667v1 Announce Type: cross Abstract: Selective inference (SI) provides statistically valid $p$-values for hypotheses selected by applying an algorithm to the data, correcting for the bias that arises when the same data are used both to select and to test a hypothesis.
By Teruyuki Katsuoka, Tomohiro Shiraishi, Shuichi Nishino, Ichiro Takeuchi
arXiv:2606. 01566v1 Announce Type: new Abstract: Small-to-medium scientific datasets place machine learning pipelines under two compounding pressures.
By Amanda S Barnard
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2609.38263v1 Announce Type: new
Abstract: Feature selection in neural networks remains a challenging problem, particularly in the presence of noisy or contaminated data. LassoNet is a recent ap...
By Daniela De Canditiis, Italia De Feis, Paola Stolfi
arXiv:2609.36396v1 Announce Type: cross
Abstract: As black-box machine learning models become increasingly common, extracting interpretations with uncertainty quantification has become a critical cha...
By Yinan Cheng, Lili Zheng
arXiv:2608.12057v4 Announce Type: replace
Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of eval...
By Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek
The paper investigates the bias in variable importance scores produced by tree-based methods, noting that continuous predictors are favored over categorical ones. It offers a theoretical explanation for this bias and proposes a straightforward fix: adding a small amount of noise to each categorical predictor. The authors validate the correction on both simulated and real-world datasets and integrate it with integrated path stability selection to achieve variable selection with false discovery control for mixed data.
By Jiahe Li, Omar Melikechi
arXiv:2608. 12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method.
By Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek