Hugging Face Trending Papers

Relevance-Aware Rule: Structural Deletion of Irrelevant Conditions in Decision Trees

Read the original on Hugging Face Trending Papers →

Decision trees generate interpretable if--then rules, yet they contain irrelevant conditions (IRCs). These IRCs arise from the structural mechanism of tree splitting and persist even in modern optimal sparse tree induction algorithms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
2d ago

Four Ways to Grow a Classifier and Why One of Them Cannot Learn

The paper investigates four ways to grow a classifier—adding a tree level, a hidden unit, a leaf split, and a statistically significant split—under a fixed protocol for tree‑structured and constructive models. It shows that the most natural method of deepening a soft decision tree by duplicating a leaf’s class distribution leaves the gradient of new gates identically zero, preventing learning, and proposes a small random perturbation as a fix. The other three growth decisions each provide a distinct benefit: fitting a new hidden unit to residual error yields a smaller network, splitting the leaf with the largest expected error adds sparsity, and requiring statistical significance before splitting adds no value and reduces accuracy.

By Cagri Temel
arXiv Machine Learning
Sep 3

RCProb: Probabilistic rule extraction from classification tree ensembles

RCProb is a probabilistic extension of rule extraction from tree ensembles that improves probability estimates by using smoothed atomic class-conditional evidence and a support‑adaptive mixture for final rule probabilities. Compared to RuleCOSI+, RCProb reduces median paired log‑loss by 71.9% for random forests and 62.5% for gradient boosting, while also decreasing the number of extracted rules by about 38% for both ensemble types. The method shows significant improvements in calibration metrics such as Confidence‑ECE and competitive native probability estimates, with further gains possible through post‑hoc calibration.

By Josue Obregon