arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2606. 31686v1 Announce Type: cross Abstract: Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret.
By Jesus S. Aguilar-Ruiz
arXiv:2606. 01111v1 Announce Type: new Abstract: Modern industrial recommender systems rely on thousands of heterogeneous features -- ranging from low-dimensional scalars (e.
By Yihong Huang, Chen Chu, Fei Chen, Yu Lin, Ruiduan Li, Zhihao Li
arXiv:2508. 14268v2 Announce Type: replace-cross Abstract: Feature selection and importance estimation in a model-agnostic setting is an ongoing challenge of significant interest.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2608. 12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method.
By Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek
arXiv:2608. 01586v1 Announce Type: cross Abstract: In recent years, numerous open-source software libraries have been developed for computing sets of features from univariate time series.
By Trent Henderson, Ben D. Fulcher
arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.
By Chenghui Zheng, Garvesh Raskutti
arXiv:2606. 07068v1 Announce Type: new Abstract: Background: Since 1990 many feature selection methods have been proposed across heterogeneous applications.
By Malick Ebiele, Malika Bendechache, Rob Brennan
arXiv:2607. 15774v1 Announce Type: cross Abstract: Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for downstream tasks remains under-explored.
By Davide Italo Serramazza, Thach Le Nguyen, Georgiana Ifrim
arXiv:2602. 02025v2 Announce Type: replace-cross Abstract: ML models critically depend on feature quality, yet in real-world settings, useful features are often distributed across multiple relational tables rather than a single dataset.
By Serafeim Papadias, Kostas Patroumpas, Dimitrios Skoutas
arXiv:2608. 04310v1 Announce Type: new Abstract: The Rashomon effect describes the phenomenon that many models can achieve nearly equivalent performance on the same learning task, with significant ramifications for robustness, feature importance, and customizability.
By Zakk Heile, Hayden McTavish, Margo Seltzer, Cynthia Rudin
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen