arXiv Machine Learning

Source-Learned Reliance for Selective Test-Time Adaptation of Multimodal Time Series

arXiv AI
Aug 26

Evaluating Deep Multivariate Imputation Models on Wearable Device Data

The paper introduces a new evaluation protocol for deep multivariate imputation models on wearable device data, addressing the issue of structured missingness where sensor features drop out together. Using a Garmin smartwatch dataset from an epilepsy patient, the authors generate realistic block-missing patterns from training data and show that matching the training protocol to this distribution reduces BRITS’ mean absolute error by 43%. They also extend BRITS with time‑of‑day encoding and compare it to linear interpolation and SAITS, finding that no single model dominates and that model rankings vary with evaluation design.

By Skye Goodman, Roussel Desmond Nzoyem, Leandro Junges, Peter Kissack, Yasser Qureshi, Amberly Brigden, Jeff Clark, Nawid Keshtmand
arXiv AI
Sep 16

SOTER: A Generative Time-Series Foundation Model for Wearable Human Physiological Signals

SOTER is a generative foundation model designed for wearable physiological time‑series data. It integrates cross‑channel coupling, spectrum‑guided expert specialization, and continuous‑time latent evolution, using a spatial feature‑aware backbone, a PSD‑guided mixture‑of‑experts layer, and a neural controlled differential equation decoder. Trained on 226 billion time points from five public datasets, SOTER outperforms baselines in zero‑shot forecasting, classification, and imputation across six benchmarks, and remains robust to additive noise.

By Fangke Chen, Sirry Chen, Wei Chen, Zhongyu Wei
arXiv Machine Learning
Sep 24

When Adaptation Hurts: Split Sensitivity and Person-Level Negative Transfer in Federated Wearable Onboarding

The paper evaluates six onboarding strategies for federated wearable models on five datasets using a leakage‑controlled protocol that fixes source checkpoints and separates calibration from evaluation. Results show that while average accuracy is high, person‑level performance can drop significantly, with some methods causing negative transfer for certain users. The study highlights that mean accuracy alone is insufficient and provides an auditable benchmark and failure map for future development.

By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
arXiv AI
Sep 25

When Temporal Perturbations Act Like Sensor Biases: Label-Free Auditing of Wearable Activity Recognizers

The paper introduces SpectrumAudit, a label‑sealed auditing method for wearable human‑activity recognition models that uses phase‑randomized full‑window stimuli to probe sensor biases. By replaying the DC component and a zero‑mean residual on held‑out subjects, the audit demonstrates significant accuracy drops across 27 victim models, with DC perturbations proving more harmful than AC in most cases. The study also shows that the audit can distinguish between persistent sensor offsets and zero‑mean variations under a fixed peak‑budget.

By Qingyu Wu, Yuan Wei, Renju Liu, Hua Cheng
arXiv Machine Learning
Sep 2

When Does Online Adaptation Pay on the Edge? A Leakage-Free Evaluation of Warmup, Learning-Rate Selection, and Resource Trade-offs for Time-Series Forecasting

The paper investigates when online adaptation benefits edge time‑series forecasting under distribution drift, using a leakage‑free streaming protocol on six public multivariate datasets. It shows that the warmup budget for static baselines and the choice of learning rate can bias perceived adaptation gains, and that a validation‑only procedure selecting warmup and optimizer rates yields Adam outperforming SGD with momentum in most settings. The study also examines accuracy versus adaptation‑state memory and per‑update latency for different adaptation strategies, highlighting parameter‑efficient variants that are nondominated on the memory axis.

By Takumi Fujimoto, Hiroaki Nishi
arXiv AI
Sep 2

Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You

The paper introduces CASTER, a gradient‑free test‑time adaptation method that keeps the model frozen by storing source class statistics in a discriminative subspace and applying an affine transformation estimated from target‑batch moments. CASTER avoids backward passes, optimizer state, and large feature banks, outperforming k‑NN on frozen features in most backbone‑dataset settings while using far less memory. The authors also propose a residual‑to‑margin transportability certificate that flags when affine transport is unreliable, and demonstrate that gating based on this certificate can recover performance losses.

By Salim Khazem, Ibrahim Mohamed Serouis