arXiv Machine Learning By Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang

Interpretable and Calibrated Classification of Clinical Data Using Supervised Feature Binarization

Read the original on arXiv Machine Learning →

The paper introduces a statistically grounded framework for interpretable, rule-based clinical classification using Bernoulli Naïve Bayes (BNB). It employs supervised chi‑square‑guided binarization to convert continuous medical variables into binary indicators, enabling BNB to handle continuous data while maintaining transparency. On three benchmark datasets—Pima Indians Diabetes, Wisconsin Breast Cancer, and Heart Failure Prediction—the method achieved AUCs of 0.800, 0.984, and 0.919, respectively, and demonstrated reliable probability calibration through cross‑validated analysis and post‑hoc beta calibration.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

Machine-Learning Assessment of the Predictive Value of Inflammatory Biomarkers for Cognitive Impairment in an Older Hispanic Adult Cohort

The study evaluates whether inflammatory biomarkers can predict cognitive impairment in older Hispanic adults using interpretable machine learning on a small clinical dataset. A leakage‑safe Bernoulli/Categorical Naive Bayes model was trained on 165 participants from the Panama Aging Research Initiative, with continuous predictors discretized via supervised chi‑square and income treated categorically. The biomarker I‑309 (CCL1) emerged as the sole reliable incremental predictor, boosting ROC‑AUC from 0.630 to 0.740 and achieving statistically significant performance across repeated cross‑validation and random partitions.

By Antony Garcia, Gabrielle Britton, Alcibiades Villarreal, Diana Oviedo, Giselle Rangel, Xinming Huang
arXiv Machine Learning
1d ago

From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning

The paper introduces a four-step pipeline that mines decision rules in the latent space of an FT-Transformer and then translates those rules back into measurable clinical features. By treating embedding dimensions that separate patient groups as latent biomarkers, small decision trees are used to extract rules, which are then mapped to raw features using gradient-input saliency and CLS attention attribution. Across six public clinical datasets, the translated rules generally outperformed raw-feature rules, achieving significant AUROC gains, though some high-performing latent rules could not be fully captured by simple raw-feature conditions.

By Majid Lotfian Delouee, Hamed Ayoobi, Sjors G. J. G. In 't Veld, Martijn C. Schut