arXiv Machine Learning

i-IF-Learn: Iterative Feature Selection and Unsupervised Learning for High-Dimensional Complex Data

arXiv:2603. 24025v2 Announce Type: replace Abstract: Unsupervised learning of high-dimensional data is challenging due to irrelevant or noisy features obscuring underlying structures.

arXiv Machine Learning
Jun 15

Cluster LOCO: Feature Importance For Interpreting Clusters

arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.

By Claire M. He, Genevera I. Allen
arXiv Machine Learning
Jul 17

Cross-Cluster Weighted Forests

arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.

By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv Machine Learning
6d ago

Towards Truly Unsupervised Evaluation of Feature Selection

arXiv:2608. 12057v1 Announce Type: new Abstract: Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method.

By Hafiz Saud Arshad, Muhammad Rajabinasab, Arthur Zimek