The paper introduces Interpretable Network-assisted Random Forest+ (RF+), a family of flexible models that combine the predictive power of random forests with network information. It offers intrinsic interpretability by providing global and local feature importance measures, as well as sample influence metrics, allowing researchers to assess both feature effects and the contribution of network neighbors. The authors claim that RF+ achieves competitive prediction accuracy while remaining transparent, making it suitable for high-impact problems where understanding model decisions is crucial.
By Tiffany M. Tang, Elizaveta Levina, Ji Zhu
arXiv:2606. 10770v1 Announce Type: cross Abstract: Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a variety of tasks such as assisting model interpretation, model selection and diagnosis, and cost-bounded learning etc.
By Guancheng Zhou, Haiping Xu, Jason Liu, Donghui Yan
arXiv:2608. 05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data.
By Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
The paper introduces Adaptive Derivative-Ordered Random Explanation (ADORE), a unified framework that uses first- and second-order derivatives to capture nonlinear feature interactions and feature-sample dynamics. ADORE combines global feature importance with local sample contributions, quantifying both magnitude and direction of feature impact while identifying critical samples. It achieves computational efficiency via randomized SVD and dynamic sparsity detection, outperforming LIME and SHAP across tabular, text, and image data, and is released as an open-source Python package on GitHub.
By Lemen Chao, Ming Lei, Anran Fanga
arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen
arXiv:2512. 13003v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is essential for determining when a supervised model encounters inputs that differ meaningfully from its training distribution.
By Min Lu, Hemant Ishwaran
arXiv:2607. 14096v1 Announce Type: new Abstract: In predictive modeling, the ability to explain why a model produces a given target prediction has become increasingly important [5, 10].
By Emiliano Massi
The paper investigates the bias in variable importance scores produced by tree-based methods, noting that continuous predictors are favored over categorical ones. It offers a theoretical explanation for this bias and proposes a straightforward fix: adding a small amount of noise to each categorical predictor. The authors validate the correction on both simulated and real-world datasets and integrate it with integrated path stability selection to achieve variable selection with false discovery control for mixed data.
By Jiahe Li, Omar Melikechi
arXiv:2609.24126v1 Announce Type: cross
Abstract: Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important fe...
By Xuhui Liu, Lili Zheng
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
The paper refactors and expands the scikit-rebate Python package, adding new Relief‑Based Algorithm (RBA) variants such as SWRF*, mu‑Relief, and five novel methods that use alternative neighbor selection and feature scoring strategies. Benchmarking across diverse genomic simulations shows that most RBAs, except mu‑Relief, effectively detect 2‑way interactions in noisy data, with far‑scoring variants like MultiSWRFDB* excelling at interaction detection but being less sensitive to main effects. The refactored package achieves 10‑ to 35‑fold runtime reductions, and the new RBAs maintain strong performance for both main effects and 2‑way epistatic interactions, preserving predictive signals for downstream modeling.
By Kia Kazemi-Nia, Harsh Bandhey, Philip J. Freda, Ryan J. Urbanowicz