arXiv Machine Learning

A Kernel Fisher Discriminant Analysis-Based Tree Ensemble Classifier: KFDA Forest

arXiv:2606. 29053v1 Announce Type: new Abstract: In general, an ensemble classifier is more accurate than a single classifier.

arXiv Machine Learning
Jun 18

Kernel of Partition Paths: A Unified Representation for Tree Ensembles

arXiv:2606. 18853v1 Announce Type: cross Abstract: A recent line of work has reframed individual decision trees as linear models on engineered features associated with their splits, opening routes for oracle inequalities and feature-importance reinterpretation, but leaving open the question of what unified geometric object a forest induces when one indexes its feature map by nodes rather than by splits.

By Nicolas Mahler
arXiv Machine Learning
Jun 30

Gradient boosting with vector-valued leafs

arXiv:2606. 29326v1 Announce Type: cross Abstract: Gradient boosting in the form of decision tree ensembles has successfully been applied to a variety of problems using simple objective functions based on log-likelihoods of a single variable.

By David Cortes
Hugging Face Trending Papers
Aug 20

DICS: Data-Informed Centroid Splitting for Decision Tree Classifiers

The paper introduces DICS, a clustering-based framework that uses data-informed priors to construct a compact set of candidate splits for decision tree classifiers. By incorporating class-aware structure, DICS reduces the split search space, preserving predictive performance while cutting training time. The authors provide theoretical analysis and experimental results showing comparable accuracy to exhaustive search across synthetic and benchmark datasets.

arXiv Machine Learning
Sep 21

Improving the Predictive Performance of Bootstrap Aggregating by Dirichlet Resampling

The paper revisits Breiman’s insight that lowering inter‑tree correlation can boost random forest performance. It introduces two new variants—Dirichlet‑Multinomial Bagging Random Forest (DM) and Dirichlet‑Weighted Random Forest (DW)—which adjust sample reweighting through a concentration parameter α>0. A theoretical criterion is presented to determine when these methods behave like standard random forests, guiding a lightweight tuning approach. Experiments on public classification benchmarks show DM and DW consistently match or outperform other random‑forest baselines with minimal extra runtime.

By Quoc Viet Le, Joonha Park
arXiv Machine Learning
Sep 3

Achieving More with Less: A Tensor-Optimization-Powered Ensemble Method

The paper proposes a tensor‑optimization‑powered ensemble method that uses confidence tensors to capture how each weak base classifier performs across different classes. By integrating these tensors and a smooth, partially convex objective that emphasizes margin, the method improves both classification accuracy and generalization while requiring fewer base learners. The authors also prove a property of the loss gradient that enables efficient gradient‑based optimization of the constrained problem.

By Jinghui Yuan, Weijin Jiang, Zhe Cao, Fangyuan Xie, Rong Wang, Feiping Nie, Yuan Yuan