BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv:2606. 00345v1 Announce Type: new Abstract: Wearable and mobile sensing technologies enable continuous monitoring of human behavior and health in real-world settings.
By Flavio Di Martino, Mattia G. Campana, Marcello Magno, Lorenza Pratali, Franca Delmastro
arXiv:2411. 15240v5 Announce Type: replace-cross Abstract: Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavioral trends.
By Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh, SouYoung Jin, Nicholas C. Jacobson
arXiv:2609.36619v1 Announce Type: new
Abstract: Polysomnography (PSG) integrates multiple physiological signals to provide a comprehensive characterization of human sleep, yet its heterogeneous chann...
By Junyu Chen, Chenxi Liu, Shiqin Tang, Hao Miao, Wanyun Ling, Ziyue Li, Hongbin Liu, Gaofeng Meng
The study introduces a clinician‑in‑the‑loop benchmark to assess whether large language models can generate evidence‑grounded Brief Hierarchical Taxonomy of Psychopathology (B‑HiTOP) item profiles from multimodal data, including passive sensing, ecological momentary assessment, and questionnaires. Using the GLOBEM dataset, the authors create 14,592 participant‑day instances aligned to 29 B‑HiTOP items across five spectra, and evaluate evidence compatibility rather than diagnostic accuracy. Two‑stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, yielding more conservative score distributions across models, spectra, and evidence settings.
By Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
arXiv:2608.28152v1 Announce Type: cross
Abstract: Agitation fluctuates over short time horizons in people living with dementia, yet continuous physiological information for anticipating next-day risk...
By Zhen Liu, Marta Bono, Robbe Decloedt, Ajda Flisar, Maarten Van Den Bossche, Maarten De Vos
arXiv:2606. 00884v1 Announce Type: cross Abstract: We study cross-subject emotion recognition from EEG, a practically important yet challenging problem in brain-computer interfaces.
By Jiaxin Qing, Lexin Li
SOTER is a generative foundation model designed for wearable physiological time‑series data. It integrates cross‑channel coupling, spectrum‑guided expert specialization, and continuous‑time latent evolution, using a spatial feature‑aware backbone, a PSD‑guided mixture‑of‑experts layer, and a neural controlled differential equation decoder. Trained on 226 billion time points from five public datasets, SOTER outperforms baselines in zero‑shot forecasting, classification, and imputation across six benchmarks, and remains robust to additive noise.
By Fangke Chen, Sirry Chen, Wei Chen, Zhongyu Wei
BioSync is a transformer-based model that fuses cardiac, neural, behavioral, and speech data from wearables and mobile devices into a continuous composite digital biomarker called the BioSync Index (BSI). The architecture uses multi-head self-attention on modality tokens and a linear branch for feature concatenation, inspired by latent-variable measurement theory. Evaluations on synthetic cohorts for cognitive decline and metabolic-autonomic conditions show BioSync achieving AUCs of 0.928 and 0.764 accuracy/F1 of 0.766, outperforming simple concatenation and other fusion strategies in most corruption scenarios.
By Seyed Mahmoud Sajjadi Mohammadabadi
arXiv:2607. 27635v1 Announce Type: cross Abstract: Wearable sensors continuously capture fine-grained multivariate time-series data, providing opportunities to model behavioural patterns associated with health outcomes.
By Xiaotong Yu, Joshua Y. Kim, HaeJin Lee, Kalina Yacef
arXiv:2606. 14604v1 Announce Type: cross Abstract: Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking.
By Pavlos Nicolaou, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios
arXiv:2607. 22508v1 Announce Type: new Abstract: Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks.
By Athanasios Papastathopoulos-Katsaros, Steven T. Lee, Lin Yao, Ajay Thomas, Junseok Park, Matthew J. McGinley, Zhandong Liu