Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary. Early-stopping strategies avoid fixing such a range, but can be sensitive to score noise and prone to premature stopping.
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2602. 05786v3 Announce Type: replace Abstract: Tree-boosting is a widely used machine learning technique for tabular data.
By Floris Jan Koster, Fabio Sigrist
arXiv:2606. 30837v1 Announce Type: cross Abstract: The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases training and prediction cost.
By Andrey A. Dukhovny, Andrey M. Lange
The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.
By Quoc Viet Le, Joonha Park
arXiv:2603.00326v2 Announce Type: replace
Abstract: Sparse oblique (SPO), part of the top-ranked configuration of Google's Yggdrasil Decision Forests (YDF), improve the accuracy while maintaining int...
By Ariel Lubonja, Jungsang Yoon, Haoyin Xu, Yue Wan, Yilin Xu, Richard Stotz, Mathieu Guillame-Bert, Joshua T. Vogelstein, Randal Burns
arXiv:2605. 20716v5 Announce Type: replace Abstract: Random forests construct each tree with a different, randomised representation of the feature space.
By Youngjoon Park
arXiv:2409. 12788v3 Announce Type: replace Abstract: Recently there has been a surge of interest in optimal decision tree (ODT) methods that globally optimize accuracy directly, in contrast to traditional approaches that locally optimize an impurity or information metric.
By Jacobus G. M. van der Linden, Dani\"el Vos, Mathijs M. de Weerdt, Sicco Verwer, Emir Demirovi\'c
arXiv:2603. 12507v2 Announce Type: replace Abstract: Minimising a spectral risk objective, defined as a weighted combination of expected cost and Conditional Value-at-Risk (CVaR), is challenging when the uncertainty distribution is decision-dependent, making both surrogate modelling and simulation-based ranking sensitive to tail estimation error.
By Marcell T. Kurbucz
arXiv:2608. 03111v1 Announce Type: new Abstract: Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width.
By Ryuichi Kanoh
CANOPY is a multi‑fidelity tree bandit algorithm that learns where a piecewise‑smooth prior holds instead of assuming global smoothness. It uses cheap random‑path probes to certify local aggregation bias and then focuses expensive leaf evaluations on cells where smoothness is violated. The method achieves provable fixed‑budget and regret guarantees that scale with the number of discontinuities, matching smooth‑tree rates when no violations exist and approaching structure‑blind search when violations are dense.
By Michael Jerge, Suman Jana
Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been established.