arXiv:2608. 03111v1 Announce Type: new Abstract: Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width.
By Ryuichi Kanoh
Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width. For gradient boosting decision trees (GBDTs), however, an analogous single-axis capacity parameter has not been established.
arXiv:2603.00326v2 Announce Type: replace
Abstract: Sparse oblique (SPO), part of the top-ranked configuration of Google's Yggdrasil Decision Forests (YDF), improve the accuracy while maintaining int...
By Ariel Lubonja, Jungsang Yoon, Haoyin Xu, Yue Wan, Yilin Xu, Richard Stotz, Mathieu Guillame-Bert, Joshua T. Vogelstein, Randal Burns
arXiv:2606. 26337v1 Announce Type: new Abstract: Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feature histograms.
By Yan Song
RCProb is a probabilistic extension of rule extraction from tree ensembles that improves probability estimates by using smoothed atomic class-conditional evidence and a support‑adaptive mixture for final rule probabilities. Compared to RuleCOSI+, RCProb reduces median paired log‑loss by 71.9% for random forests and 62.5% for gradient boosting, while also decreasing the number of extracted rules by about 38% for both ensemble types. The method shows significant improvements in calibration metrics such as Confidence‑ECE and competitive native probability estimates, with further gains possible through post‑hoc calibration.
By Josue Obregon
Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feature histograms. Existing approaches such as random feature subsampling (feature_fraction) discard features without regard for their predictive utility.