arXiv:2512. 11081v2 Announce Type: replace-cross Abstract: Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models.
By Kata Vuk, Nicolas Alexander Ihlo, Merle Behr
The paper investigates the bias in variable importance scores produced by tree-based methods, noting that continuous predictors are favored over categorical ones. It offers a theoretical explanation for this bias and proposes a straightforward fix: adding a small amount of noise to each categorical predictor. The authors validate the correction on both simulated and real-world datasets and integrate it with integrated path stability selection to achieve variable selection with false discovery control for mixed data.
By Jiahe Li, Omar Melikechi
The paper introduces Interpretable Network-assisted Random Forest+ (RF+), a family of flexible models that combine the predictive power of random forests with network information. It offers intrinsic interpretability by providing global and local feature importance measures, as well as sample influence metrics, allowing researchers to assess both feature effects and the contribution of network neighbors. The authors claim that RF+ achieves competitive prediction accuracy while remaining transparent, making it suitable for high-impact problems where understanding model decisions is crucial.
By Tiffany M. Tang, Elizaveta Levina, Ji Zhu
arXiv:2608. 05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data.
By Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal