Hugging Face Trending Papers

EMA-FS: Accelerating GBDT Training via Gain-Informed Feature Screening

Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feature histograms. Existing approaches such as random feature subsampling (feature_fraction) discard features without regard for their predictive utility.

arXiv Machine Learning
2d ago

Vectorized Dynamic Histograms for Sparse Oblique Forests

arXiv:2603.00326v2 Announce Type: replace Abstract: Sparse oblique (SPO), part of the top-ranked configuration of Google's Yggdrasil Decision Forests (YDF), improve the accuracy while maintaining int...

By Ariel Lubonja, Jungsang Yoon, Haoyin Xu, Yue Wan, Yilin Xu, Richard Stotz, Mathieu Guillame-Bert, Joshua T. Vogelstein, Randal Burns
arXiv Machine Learning
Sep 3

RCProb: Probabilistic rule extraction from classification tree ensembles

RCProb is a probabilistic extension of rule extraction from tree ensembles that improves probability estimates by using smoothed atomic class-conditional evidence and a support‑adaptive mixture for final rule probabilities. Compared to RuleCOSI+, RCProb reduces median paired log‑loss by 71.9% for random forests and 62.5% for gradient boosting, while also decreasing the number of extracted rules by about 38% for both ensemble types. The method shows significant improvements in calibration metrics such as Confidence‑ECE and competitive native probability estimates, with further gains possible through post‑hoc calibration.

By Josue Obregon