arXiv Machine Learning

Distributional Split Criteria for Random Forests: Extensions, Shrinkage, and the Robustness of Mean Splitting

arXiv:2607. 23721v1 Announce Type: cross Abstract: Distributional random forests replace mean-based CART splitting with criteria that compare the full conditional response distribution in candidate children.

arXiv Machine Learning
Jun 8

Adaptive Conditional Forest Sampling for Spectral Risk Optimisation under Decision-Dependent Uncertainty

arXiv:2603. 12507v2 Announce Type: replace Abstract: Minimising a spectral risk objective, defined as a weighted combination of expected cost and Conditional Value-at-Risk (CVaR), is challenging when the uncertainty distribution is decision-dependent, making both surrogate modelling and simulation-based ranking sensitive to tail estimation error.

By Marcell T. Kurbucz
arXiv Machine Learning
Jun 18

Kernel of Partition Paths: A Unified Representation for Tree Ensembles

arXiv:2606. 18853v1 Announce Type: cross Abstract: A recent line of work has reframed individual decision trees as linear models on engineered features associated with their splits, opening routes for oracle inequalities and feature-importance reinterpretation, but leaving open the question of what unified geometric object a forest induces when one indexes its feature map by nodes rather than by splits.

By Nicolas Mahler
arXiv Machine Learning
Jul 3

Conditional Inference Trees and Forests for Feature Selection

arXiv:2607. 01417v1 Announce Type: new Abstract: Conditional inference trees (CIT) and conditional inference forests (CIF) reduce split-selection bias by testing features before choosing split thresholds, but repeated permutation tests and threshold searches can make these methods computationally expensive.

By Robert Milletich, Justin Downes, Steve Goley, Newel Hirst
arXiv Machine Learning
1d ago

MECHVAR: Variance-Guided Mechanism Discrimination for Autonomous Machine Learning Experiment Selection

MECHVAR is a lightweight, auditable rule for selecting experiments from a finite library to discriminate between candidate mechanisms. It chooses probes by maximizing the posterior‑weighted variance of predicted responses, a score that aligns with the Box–Hill pairwise‑KL criterion and links to expected information gain when separations are small. Experiments on a 25‑block audit and a Digits loop show MECHVAR outperforming confirmation‑first strategies and matching or exceeding EIG in identification accuracy while being far faster to compute.

By Yifan Guo
Hugging Face Trending Papers
Jun 2

How Many Trees in a Random Forest? A Revisited Approach with Plateau Search and Optuna Integration

Hyperparameter optimization (HPO) for Random Forest faces a specific difficulty in tuning the number of trees: the predictive score typically improves monotonically with ensemble size, so standard methods such as Tree-structured Parzen Estimator (TPE) and Hyperband require a predefined search range and often drive the estimate toward its right boundary. Early-stopping strategies avoid fixing such a range, but can be sensitive to score noise and prone to premature stopping.