arXiv Machine Learning

Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling

The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.

arXiv Machine Learning
Jun 3

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

arXiv:2606. 03549v1 Announce Type: new Abstract: Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary.

By Vadim Porvatov, Andrey Dukhovny, Andrey Lange
Hugging Face Trending Papers
Jun 2

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary. Early-stopping strategies avoid fixing such a range, but can be sensitive to score noise and prone to premature stopping.

Hugging Face Trending Papers
Jul 29

Feature Bagging Provides Stability

We study feature bagging through the lens of algorithmic stability. Feature bagging is an ensemble strategy that aggregates base learners trained on randomly subsampled feature subsets, possibly in a data-dependent manner.

arXiv Machine Learning
Jul 3

Conditional Inference Trees and Forests for Feature Selection

arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.

By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv Machine Learning
Jul 17

Cross-Cluster Weighted Forests

arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.

By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv Machine Learning
Jun 10

Correcting Variable Importance Scored by Random Forests

arXiv:2606. 10770v1 Announce Type: cross Abstract: Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a variety of tasks such as assisting model interpretation, model selection and diagnosis, and cost-bounded learning etc.

By Guancheng Zhou, Haiping Xu, Jason Liu, Donghui Yan