BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv:2606. 00345v1 Announce Type: new Abstract: Wearable and mobile sensing technologies enable continuous monitoring of human behavior and health in real-world settings.
By Flavio Di Martino, Mattia G. Campana, Marcello Magno, Lorenza Pratali, Franca Delmastro
arXiv:2411. 15240v5 Announce Type: replace-cross Abstract: Wearable movement data is collected by nearly all commercially available smartwatches and is a valuable resource for mental health research, reflecting fine-grained temporal behavioral trends.
By Franklin Y. Ruan, Aiwei Zhang, Jenny Y. Oh, SouYoung Jin, Nicholas C. Jacobson
arXiv:2609.36619v1 Announce Type: new
Abstract: Polysomnography (PSG) integrates multiple physiological signals to provide a comprehensive characterization of human sleep, yet its heterogeneous chann...
By Junyu Chen, Chenxi Liu, Shiqin Tang, Hao Miao, Wanyun Ling, Ziyue Li, Hongbin Liu, Gaofeng Meng
The study introduces a clinician‑in‑the‑loop benchmark to assess whether large language models can generate evidence‑grounded Brief Hierarchical Taxonomy of Psychopathology (B‑HiTOP) item profiles from multimodal data, including passive sensing, ecological momentary assessment, and questionnaires. Using the GLOBEM dataset, the authors create 14,592 participant‑day instances aligned to 29 B‑HiTOP items across five spectra, and evaluate evidence compatibility rather than diagnostic accuracy. Two‑stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, yielding more conservative score distributions across models, spectra, and evidence settings.
By Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
arXiv:2608.28152v1 Announce Type: cross
Abstract: Agitation fluctuates over short time horizons in people living with dementia, yet continuous physiological information for anticipating next-day risk...
By Zhen Liu, Marta Bono, Robbe Decloedt, Ajda Flisar, Maarten Van Den Bossche, Maarten De Vos