The paper introduces the Individual Conformal Coupling Monitor (ICCM), a lightweight pre‑inference tool that detects structural ambiguity—when physiological signals that appear plausible individually form a pattern poorly supported by a person’s non‑stress baseline—in wearable stress classifiers. On the WESAD dataset, a Random Forest achieves high mean accuracy but fails entirely for Subject 14 due to weakened cross‑signal coupling near stress onset. ICCM quantifies subject‑specific coupling divergence and can route data to classify, defer, or abstain, reducing false positives slightly and withholding some misclassified windows, though it does not fully correct the failure.
By Saba A. Farahani, Hung Cao, Amir M. Rahmani
The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.
By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
The paper introduces SpectrumAudit, a label‑sealed auditing method for wearable human‑activity recognition models that uses phase‑randomized full‑window stimuli to probe sensor biases. By replaying the DC component and a zero‑mean residual on held‑out subjects, the audit demonstrates significant accuracy drops across 27 victim models, with DC perturbations proving more harmful than AC in most cases. The study also shows that the audit can distinguish between persistent sensor offsets and zero‑mean variations under a fixed peak‑budget.
By Qingyu Wu, Yuan Wei, Renju Liu, Hua Cheng
arXiv:2607. 18279v1 Announce Type: cross Abstract: Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal.
By Filippo Cenacchi, Longbing Cao, Runze Yang
arXiv:2605. 22774v3 Announce Type: replace-cross Abstract: Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects.
By Amir Mousavi, Erfan Nourbakhsh, Mohammad Sadegh Sirjani, Mimi Xie, Rocky Slavin, Leslie Neely, John Davis, John Quarles
arXiv:2608. 08244v1 Announce Type: new Abstract: General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health tasks.
By Yifan Wang, Chenzhong Li
FemWear is a parameter‑efficient wearable foundation model specifically tailored for women's health. It repurposes a pretrained multimodal wearable backbone by training only 239,236 encoder parameters—just 1.11% of the original 21.54M—using low‑rank residual adapters and causal task‑family heads to create a shared longitudinal representation for menstrual, symptom, affective, sleep/recovery, autonomic, activity, and pregnancy outcomes. Evaluations across six cohorts and 63 metrics show improvements in cycle‑phase macro‑F1 and reductions in mean absolute error for cramps, mood symptoms, and sleep problems, while maintaining the OpenMHC ability‑retention benchmark.
By Yifan Wang, Chenzhong Li
arXiv:2609.22619v1 Announce Type: new
Abstract: Tracking recovery of walking function requires detecting meaningful gait change across rehabilitation sessions, yet objective 3D measurement remains co...
By Nethmi Jayasinghe, Mihir Parashar, Amit Ranjan Trivedi
arXiv:2411. 15240v5 Announce Type: replace-cross Abstract: Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavioral trends.
By Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh, SouYoung Jin, Nicholas C. Jacobson
arXiv:2608. 16087v1 Announce Type: cross Abstract: Removing wearables from physiological monitoring also removes their supervision: the signal indicating where and when a stress response occurred.
By Sachin Deb, Harshit Sharma, Asif Salekin
arXiv:2607. 20237v1 Announce Type: new Abstract: Rehabilitation scoring systems are most useful when their outputs can be reviewed and interpreted within clinical workflows.
By Yankai Zheng, Yuhe Liu, Yuxin Ma, Tianci Xue, Jiayuan Tian, Yu Fu, Yuxuan Hu, Jianing Wang, Zichun Xiao, Junya Mu, Shaohui Ma
WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.
By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda