arXiv AI

An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness

arXiv:2604. 23954v2 Announce Type: replace Abstract: Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decision-making.

arXiv AI
Aug 20

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.

By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv AI
Jul 21

Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis

arXiv:2607. 16253v1 Announce Type: cross Abstract: Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world applications due to deficient external testing and fairness assessment.

By Rajveer Singh Pall, Sameer Yadav, Siddharth Bhalerao, Sourabh Sahu, Ritu Ahluwalia, Bhaskar Awadhiya
arXiv Machine Learning
Jun 8

GlucoFM-Bench: Benchmarking Time-Series Foundation Models for Blood Glucose Forecasting

arXiv:2606. 06881v1 Announce Type: new Abstract: Blood glucose forecasting models are foundational for modern diabetes management systems, as reliable short-term predictions can enable proactive interventions, support automated insulin delivery, and reduce the risk of hypo- and hyperglycemic events.

By Baiying Lu, Zhaohui Liang, Ryan Pontius, Shengpu Tang, Temiloluwa Prioleau
arXiv Machine Learning
Sep 18

Interpretable and Calibrated Classification of Clinical Data Using Supervised Feature Binarization

The paper introduces a statistically grounded framework for interpretable, rule-based clinical classification using Bernoulli Naïve Bayes (BNB). It employs supervised chi‑square‑guided binarization to convert continuous medical variables into binary indicators, enabling BNB to handle continuous data while maintaining transparency. On three benchmark datasets—Pima Indians Diabetes, Wisconsin Breast Cancer, and Heart Failure Prediction—the method achieved AUCs of 0.800, 0.984, and 0.919, respectively, and demonstrated reliable probability calibration through cross‑validated analysis and post‑hoc beta calibration.

By Antony Garcia, Adrian Noriega, Gabrielle Britton, Xinming Huang
arXiv Statistics ML
Aug 25

Primal--Dual Alternating Neural Learning for Timely Classification with Performance Guarantees

The paper introduces a new method for timely risk classification in clinical monitoring, framing the problem as a multi‑objective optimization that balances early classification, sensitivity, specificity, and monitoring cost. It derives an optimal decision rule via a value recursion and estimates it from data using a recurrent neural network combined with a primal–dual updating scheme to enforce performance constraints. Experiments, including a case study on continuous glucose monitoring for hypoglycemia prediction, show that the approach produces accurate, timely decision rules that meet the specified operating characteristics.

By Jiaming Qiu, Yingye Zheng, Ying-Qi Zhao
arXiv Computation and Language
Sep 1

DIASENTINEL: An Auditable Multi-Agent System for Guideline-Grounded Diabetes Risk Screening

arXiv:2608.31128v1 Announce Type: new Abstract: Large language models (LLMs) offer promising clinical decision support but remain vulnerable to hallucinated facts, unsupported recommendations, and ci...

By Yung Wei Shueh, Zhi-Jie Chen, Chia-Hsuan Hsu, Hsin-Ling Hsu, Donghua Zhang, Chenwei Wu, Jun-En Ding, Tongze Zhang, Shihao Yang, Pengfei Hu, Fang-Ming Hung, Feng Liu