arXiv:2606. 10770v1 Announce Type: cross Abstract: Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a variety of tasks such as assisting model interpretation, model selection and diagnosis, and cost-bounded learning etc.
By Guancheng Zhou, Haiping Xu, Jason Liu, Donghui Yan
arXiv:2512. 11081v2 Announce Type: replace-cross Abstract: Feature and Interaction Importance (FII) methods are essential in supervised learning for assessing the relevance of input variables and their interactions in complex prediction models.
By Kata Vuk, Nicolas Alexander Ihlo, Merle Behr
arXiv:2511. 20851v3 Announce Type: replace-cross Abstract: Feature selection remains difficult in modern high-dimensional settings, and established methods such as Boruta and Recursive Feature Elimination are either computationally costly or lack a statistically justified stopping criterion for their importance scores.
By Mousam Sinha, Tirtha Sarathi Ghosh, Koushik Biswas, Ridam Pal
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2609.24528v1 Announce Type: cross
Abstract: Random forests (RFs) predict well but are opaque, whereas single decision trees are interpretable but unstable. Artificial representative trees (ARTs...
By Lea L. Mairh\"ofer, Silke Szymczak, Bj\"orn-Hergen Laabs, Tuwe L\"ofstr\"om-Cavallin
arXiv:2605. 20716v5 Announce Type: replace Abstract: Random forests construct each tree with a different, randomised representation of the feature space.
By Youngjoon Park
The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.
By Quoc Viet Le, Joonha Park
Random forests (RFs) predict well but are opaque, whereas single decision trees are interpretable but unstable. Artificial representative trees (ARTs) were developed as interpretable surrogate models...
arXiv:2602. 07453v2 Announce Type: replace Abstract: Decision tree ensembles are widely used in critical domains, making robustness and sensitivity analysis essential to their trustworthiness.
By Namrita Varshney, Ashutosh Gupta, Arhaan Ahmad, Tanay V. Tayal, S. Akshay
arXiv:2605. 13830v2 Announce Type: replace-cross Abstract: Decision tree ensembles (DTE) are a popular model for a wide range of AI classification tasks, used in multiple safety critical domains, and hence verifying properties on these models has been an active topic of study over the last decade.
By Ajinkya Naik, Chaitanya Garg, S. Akshay, Ashutosh Gupta, Kuldeep S. Meel
arXiv:2411. 08821v4 Announce Type: replace-cross Abstract: Global variable importance measures are commonly used to interpret the results of machine learning models.
By Kelvyn K. Bladen, Adele Cutler, D. Richard Cutler, Kevin R. Moon
The paper introduces Interpretable Network-assisted Random Forest+ (RF+), a family of flexible models that combine the predictive power of random forests with network information. It offers intrinsic interpretability by providing global and local feature importance measures, as well as sample influence metrics, allowing researchers to assess both feature effects and the contribution of network neighbors. The authors claim that RF+ achieves competitive prediction accuracy while remaining transparent, making it suitable for high-impact problems where understanding model decisions is crucial.
By Tiffany M. Tang, Elizaveta Levina, Ji Zhu