Non-invasive blood glucose level (BGL) estimation from photoplethysmography (PPG) holds great promise for wearable health monitoring, but results across studies are hard to compare due to inconsistent datasets, data leakage, and non-standardized evaluation metrics. We present the first reproducible, extensible evaluation pipeline and use it to reassess five representative PPG-based BGL methods on published datasets under three increasingly strict data-split protocols: random window-level, participant-aware, and leave-some-participants-out (LSPO).
FairGlucose is a 300‑patient CGM cohort balanced across 12 demographic strata, providing 132,480 forecasting samples and 3,945 behavioral events. Benchmarking 33 models on 2‑hour glucose forecasting revealed that population‑level validation masks significant subgroup disparities, with error ratios ranging from 0.8 to 1.4 and T1D patients experiencing 6 mg/dL higher error than T2D. The study shows that these gaps persist across all models, align with clinically hard cases, and vary with input‑length sensitivity, underscoring the need for subgroup‑disaggregated reporting in digital health AI.
By Junjie Luo, Xuzhe Zhi, Rui Han, Abhimanyu Kumbara, Anand K. Iyer, Mansur E. Shomali, Ritu Agarwal, Guodong Gordon Gao
arXiv:2608. 00943v1 Announce Type: cross Abstract: Automated sleep staging assigns discrete stage labels to successive time epochs throughout an overnight recording; conventionally each window spans at least 30 seconds, reflecting the minimum temporal resolution of the clinical scoring standard.
By Shuntian Zheng, Jiawei Wang, Cong Fu, Huan Yu, Chen Chen, Yu Guan, Sai Gu
arXiv:2607. 21117v1 Announce Type: cross Abstract: Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes management, particularly for type 1 diabetes.
By Davide Marelli, Giorgia Rigamonti, Mirko Paolo Barbato, Paolo Napoletano
arXiv:2606. 16056v1 Announce Type: new Abstract: Dysglycemia, encompassing both prediabetes and diabetes, affects huge numbers of adults worldwide, yet many of them remain undiagnosed.
By Black Sun, Chenyi Zhang, Kaiyi Ji, Xi Lu
arXiv:2606. 12699v1 Announce Type: cross Abstract: Type 2 Diabetes (T2D) poses an increasing global health threat, demanding effective glycemic assessment to support personalized and improved diabetes care.
By Yifan Gao, Yanmin Gong, Yun Shi, Yuanxiong Guo
arXiv:2609.08772v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only...
By Andrea Apicella, Pasquale Arpaia, Matteo Orefice, Andrea Pollastro, Roberto Prevete
GlucoFM is a lightweight foundation model for continuous glucose monitoring that aligns irregular CGM data to a 24‑hour grid and splits glucose dynamics into slow‑varying trend and short‑term deviation streams. Pre‑trained on over 109,000 hours of unlabeled recordings, it outperforms existing CGM‑specific models on seven phenotype‑classification tasks, improving average PR‑AUC by 4.1 points and enabling strong cross‑dataset transfer and few‑shot adaptation. When combined with meal, nutrition, and subject context, its frozen encoder delivers the lowest two‑hour postprandial glycemic response errors for trajectory, incremental AUC, peak rise, and peak timing metrics.
By Zechen Li, Keerthana Natarajan, Weizhi Zhang, Menglian Zhou, Simon A. Lee, Yuwei Zhang, Maxwell A. Xu, Zeinab Esmaeilpour, Flora D. Salim, Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, Ahmed A. Metwally
The paper examines how the representation of physiological data affects the performance of large language models (LLMs) in predicting post‑meal blood glucose events for people with type 1 diabetes. Using the OhioT1DM dataset, the authors compare zero‑shot and few‑shot prompt‑based LLMs across 30, 60, and 90‑minute horizons, varying the textual encoding of glucose readings, derived descriptors, and contextual variables such as insulin, meals, carbs, and activity. Results show that while conventional supervised models excel at hyperglycemia prediction, certain prompt‑based LLM configurations outperform them for hypoglycemia, and that the way data is presented to the model is a key determinant of success, with added context not consistently improving outcomes.
The study evaluates time‑series foundation models for continuous glucose monitoring (CGM) forecasting across eight public datasets covering Type 1, Type 2, and non‑diabetes populations. Zero‑shot foundation models did not consistently beat strong task‑specific baselines, but lightweight fine‑tuning of models like Chronos‑Bolt improved root‑mean‑square error by up to 18% in both in‑distribution and out‑of‑distribution settings. Incorporating multimodal dietary context via CGMacros and a residual‑based fusion framework further reduced overall RMSE by ~3% and postprandial RMSE by ~15%, indicating that dietary signals add clinically meaningful value beyond CGM alone.
By Bowen Zhang, Hsiu-Wen Cheng, Hongyu Yang, Evie L. Shen, Joleen Vansomphone, Yuna Li, Kerry Zhou, Zitian Qu, Suning Zhao, Xiangning Deng, Hua Zhou, Jin J. Zhou
arXiv:2606. 06881v1 Announce Type: new Abstract: Blood glucose forecasting models are foundational for modern diabetes management systems, as reliable short-term predictions can enable proactive interventions, support automated insulin delivery, and reduce the risk of hypo- and hyperglycemic events.
By Baiying Lu, Zhaohui Liang, Ryan Pontius, Shengpu Tang, Temiloluwa Prioleau
arXiv:2606. 15927v1 Announce Type: new Abstract: Diabetes and extreme blood sugar levels are some of the major health problems faced by humans today across the world.
By Ruhani Bhatia, Vijval Ekbote