FemWear is a parameter‑efficient wearable foundation model specifically tailored for women's health. It repurposes a pretrained multimodal wearable backbone by training only 239,236 encoder parameters—just 1.11% of the original 21.54M—using low‑rank residual adapters and causal task‑family heads to create a shared longitudinal representation for menstrual, symptom, affective, sleep/recovery, autonomic, activity, and pregnancy outcomes. Evaluations across six cohorts and 63 metrics show improvements in cycle‑phase macro‑F1 and reductions in mean absolute error for cramps, mood symptoms, and sleep problems, while maintaining the OpenMHC ability‑retention benchmark.
By Yifan Wang, Chenzhong Li
arXiv:2411. 15240v5 Announce Type: replace-cross Abstract: Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavioral trends.
By Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh, SouYoung Jin, Nicholas C. Jacobson
SleepFM-2 is a foundation model trained on 282,511 polysomnography recordings, covering over two million hours of multimodal sleep physiology. It outperforms its predecessor in disease prediction, sleep scoring, and event detection, and its representation improves performance across diverse tasks—including wearable sensing, subjective sleep reports, and transfer to other EEG modalities. When combined with age, sex, and BMI, the model meets stringent discrimination criteria for 215 EHR phenotypes, adding reproducible information beyond demographics for 155 of them.
By Rahul Thapa, Christopher Sun, William Theodor Lehn-Schioler, Sophia Claire Kivelson, Umaer Hanif, Hyatt Moore IV, Harrison G. Zhang, Hafsa Ahmed, Marcus Dige, Niels R. Lorenzen, Elisabeth Roxane M. Heremans, Adrien Specht, Ulysse Gimenez, Robin Guillard, Andreas Brink-Kjaer, James Zou, Emmanuel Mignot
arXiv:2606. 07692v1 Announce Type: cross Abstract: Foundation models for wearable biosignals have matched or exceeded supervised specialists across a range of clinical tasks, yet all rely on modalities that require deliberate user action--wearing a device or visiting a sleep lab.
By Magnus Ruud Kjaer, Haejun Han, Ashish Neupane, David Q. Sun
arXiv:2607. 15721v1 Announce Type: new Abstract: Cardiometabolic diseases remain among the most persistent drivers of preventable morbidity because diabetes, hypertension, and cardiovascular disease frequently co-occur and share metabolic, vascular, demographic, and behavioral determinants.
By S M Asif Hossain, Ruksat Khan Shayoni, M. F. Mridha, Jungpil Shin
arXiv:2606. 18506v1 Announce Type: new Abstract: Objective sleep assessment relies on polysomnography (PSG), yet clinical impact is often better reflected in patient-reported outcomes (PROs) such as sleepiness and fatigue.
By Saba A. Farahani, Elahe Khatibi, Manoj Vishwanath, Amir M. Rahmani, Hung Cao
The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.
By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
BioSync is a transformer-based model that fuses cardiac, neural, behavioral, and speech data from wearables and mobile devices into a continuous composite digital biomarker called the BioSync Index (BSI). The architecture uses multi-head self-attention on modality tokens and a linear branch for feature concatenation, inspired by latent-variable measurement theory. Evaluations on synthetic cohorts for cognitive decline and metabolic-autonomic conditions show BioSync achieving AUCs of 0.928 and 0.764 accuracy/F1 of 0.766, outperforming simple concatenation and other fusion strategies in most corruption scenarios.
By Seyed Mahmoud Sajjadi Mohammadabadi
arXiv:2609.06080v1 Announce Type: cross
Abstract: Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can r...
By Gal Sapir, Alon Diament, Adva Wolf, Doron Yaya-Stupp, Dikla Gelbard Solodkin, Dana Azouri, Anat Etzion-Fuchs, Guy Lutsker, Eran Segal, Hagai Rossman
The study evaluates time‑series foundation models for continuous glucose monitoring (CGM) forecasting across eight public datasets covering Type 1, Type 2, and non‑diabetes populations. Zero‑shot foundation models did not consistently beat strong task‑specific baselines, but lightweight fine‑tuning of models like Chronos‑Bolt improved root‑mean‑square error by up to 18% in both in‑distribution and out‑of‑distribution settings. Incorporating multimodal dietary context via CGMacros and a residual‑based fusion framework further reduced overall RMSE by ~3% and postprandial RMSE by ~15%, indicating that dietary signals add clinically meaningful value beyond CGM alone.
By Bowen Zhang, Hsiu-Wen Cheng, Hongyu Yang, Evie L. Shen, Joleen Vansomphone, Yuna Li, Kerry Zhou, Zitian Qu, Suning Zhao, Xiangning Deng, Hua Zhou, Jin J. Zhou
GlucoFM is a lightweight foundation model for continuous glucose monitoring that aligns irregular CGM data to a 24‑hour grid and splits glucose dynamics into slow‑varying trend and short‑term deviation streams. Pre‑trained on over 109,000 hours of unlabeled recordings, it outperforms existing CGM‑specific models on seven phenotype‑classification tasks, improving average PR‑AUC by 4.1 points and enabling strong cross‑dataset transfer and few‑shot adaptation. When combined with meal, nutrition, and subject context, its frozen encoder delivers the lowest two‑hour postprandial glycemic response errors for trajectory, incremental AUC, peak rise, and peak timing metrics.
By Zechen Li, Keerthana Natarajan, Weizhi Zhang, Menglian Zhou, Simon A. Lee, Yuwei Zhang, Maxwell A. Xu, Zeinab Esmaeilpour, Flora D. Salim, Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, Ahmed A. Metwally
The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.
By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan