arXiv AI
Aug 26

Evaluating Deep Multivariate Imputation Models on Wearable Device Data

The paper introduces a new evaluation protocol for deep multivariate imputation models on wearable device data, addressing the issue of structured missingness where sensor features drop out together. Using a Garmin smartwatch dataset from an epilepsy patient, the authors generate realistic block-missing patterns from training data and show that matching the training protocol to this distribution reduces BRITS’ mean absolute error by 43%. They also extend BRITS with time‑of‑day encoding and compare it to linear interpolation and SAITS, finding that no single model dominates and that model rankings vary with evaluation design.

By Skye Goodman, Roussel Desmond Nzoyem, Leandro Junges, Peter Kissack, Yasser Qureshi, Amberly Brigden, Jeff Clark, Nawid Keshtmand
arXiv AI
Sep 16

SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals

SOTER is a generative foundation model designed for wearable physiological time‑series data. It integrates cross‑channel coupling, spectrum‑guided expert specialization, and continuous‑time latent evolution, using a spatial feature‑aware backbone, a PSD‑guided mixture‑of‑experts layer, and a neural controlled differential equation decoder. Trained on 226 billion time points from five public datasets, SOTER outperforms baselines in zero‑shot forecasting, classification, and imputation across six benchmarks, and remains robust to additive noise.

By Fangke Chen, Sirry Chen, Wei Chen, Zhongyu Wei
arXiv Machine Learning
Sep 24

When Adaptation Hurts: Split Sensitivity and Person-Level Negative Transfer in Federated Wearable Onboarding

The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.

By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta