arXiv:2609.13192v1 Announce Type: new
Abstract: This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using...
By Feisal Alaswad, Batoul Aljaddouh, Maher Alrahhal, Wafaa Al Nassan, Talal Bonn
The paper presents optimizations for Sparse Oblique (SPO) forests in Google’s Yggdrasil Decision Forests, addressing training speed issues caused by runtime sampling of sparse linear feature combinations. By fixing inefficiencies and introducing two new methods—hierarchical AVX2/AVX-512 vectorized histogram filling and runtime‑dynamic histograms—the authors achieve 2–5× speedups for both Gradient Boosted Trees and Random Forests, bringing SPO‑RF training time on par with axis‑aligned RFs. Extensive evaluation on 19 datasets, including up to 10.5 million rows and 1.6 million features, demonstrates these improvements without compromising accuracy.
By Ariel Lubonja, Jungsang Yoon, Haoyin Xu, Yue Wan, Yilin Xu, Richard Stotz, Mathieu Guillame-Bert, Joshua T. Vogelstein, Randal Burns
arXiv:2606. 26337v1 Announce Type: new Abstract: Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feature histograms.
By Yan Song
arXiv:2610.02610v1 Announce Type: new
Abstract: The research presented here focuses on modeling machine-learning performance. The thesis introduces Seer, a system that generates empirical observation...
By Carl Myers Kadie
arXiv:2608. 03111v1 Announce Type: new Abstract: Double descent is commonly studied by scaling an explicit capacity parameter, such as neural-network width.
By Ryuichi Kanoh
Gradient Boosted Decision Trees (GBDT), exemplified by LightGBM, spend a dominant fraction of training time -- typically 65-70% -- constructing per-feature histograms. Existing approaches such as random feature subsampling (feature_fraction) discard features without regard for their predictive utility.