arXiv:2512. 11081v2 Announce Type: replace-cross Abstract: Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models.
By Kata Vuk, Nicolas Alexander Ihlo, Merle Behr
The paper investigates the bias in variable importance scores produced by tree-based methods, noting that continuous predictors are favored over categorical ones. It offers a theoretical explanation for this bias and proposes a straightforward fix: adding a small amount of noise to each categorical predictor. The authors validate the correction on both simulated and real-world datasets and integrate it with integrated path stability selection to achieve variable selection with false discovery control for mixed data.
By Jiahe Li, Omar Melikechi
The paper introduces Interpretable Network-assisted Random Forest+ (RF+), a family of flexible models that combine the predictive power of random forests with network information. It offers intrinsic interpretability by providing global and local feature importance measures, as well as sample influence metrics, allowing researchers to assess both feature effects and the contribution of network neighbors. The authors claim that RF+ achieves competitive prediction accuracy while remaining transparent, making it suitable for high-impact problems where understanding model decisions is crucial.
By Tiffany M. Tang, Elizaveta Levina, Ji Zhu
arXiv:2608. 05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data.
By Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.
By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.
By Quoc Viet Le, Joonha Park
arXiv:2512. 13003v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is essential for determining when a supervised model encounters inputs that differ meaningfully from its training distribution.
By Min Lu, Hemant Ishwaran
arXiv:2609.08554v1 Announce Type: new
Abstract: In data-driven training, multivariate time-series forecasting is usually optimized with a scalar loss averaged over samples, variables, and horizons. T...
By Jinwoo Park, Hyeongwon Kang, Pilsung Kang
arXiv:2609.24126v1 Announce Type: cross
Abstract: Black-box machine learning models increasingly deliver strong predictions, but extracting useful information from them, such as a set of important fe...
By Xuhui Liu, Lili Zheng
arXiv:2606. 14592v1 Announce Type: cross Abstract: Clustering is widely used for exploratory analysis and scientific discovery, driving insights from market segmentation to biological data analysis, but its outputs can be difficult to interpret, audit, and reproduce as modern datasets become increasingly large and complex.
By Claire M. He, Genevera I. Allen