RCProb is a probabilistic extension of rule extraction from tree ensembles that improves probability estimates by using smoothed atomic class-conditional evidence and a support‑adaptive mixture for final rule probabilities. Compared to RuleCOSI+, RCProb reduces median paired log‑loss by 71.9% for random forests and 62.5% for gradient boosting, while also decreasing the number of extracted rules by about 38% for both ensemble types. The method shows significant improvements in calibration metrics such as Confidence‑ECE and competitive native probability estimates, with further gains possible through post‑hoc calibration.
By Josue Obregon
arXiv:2609.24528v1 Announce Type: cross
Abstract: Random forests (RFs) predict well but are opaque, whereas single decision trees are interpretable but unstable. Artificial representative trees (ARTs...
By Lea L. Mairh\"ofer, Silke Szymczak, Bj\"orn-Hergen Laabs, Tuwe L\"ofstr\"om-Cavallin
arXiv:2609.26839v1 Announce Type: cross
Abstract: Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deploymen...
By Zeming Liu, Hang Lyu, Jingtao Zhang, Yuan Xie
Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary. Early-stopping strategies avoid fixing such a range, but can be sensitive to score noise and prone to premature stopping.
arXiv:2606. 30995v1 Announce Type: new Abstract: Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains.
By Zakk Heile, Hayden McTavish, Margo Seltzer, Cynthia Rudin
Random forests (RFs) predict well but are opaque, whereas single decision trees are interpretable but unstable. Artificial representative trees (ARTs) were developed as interpretable surrogate models...
arXiv:2606. 03549v1 Announce Type: new Abstract: Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary.
By Vadim Porvatov, Andrey Dukhovny, Andrey Lange
arXiv:2412. 16209v5 Announce Type: replace Abstract: When using machine learning for imbalanced binary classification problems, it is common to subsample the majority class to create a (more) balanced training dataset.
By Nathan Phelps, Daniel J. Lizotte, Douglas G. Woolford
The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.
By Quoc Viet Le, Joonha Park
arXiv:2605. 22740v2 Announce Type: replace Abstract: Decision trees assign identical confidence to instances near and far from each split threshold.
By William Smits
arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.
By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv:2608. 16147v1 Announce Type: new Abstract: Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were properties of the method.
By Diyorbek Musaev