arXiv AI

Mobile Interaction for Assessing Fatigue, Sleep, and Activity in Neurodegenerative and Chronic Diseases

arXiv:2608. 06380v1 Announce Type: cross Abstract: Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID).

arXiv AI
Aug 10

PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship

arXiv:2605. 17679v2 Announce Type: replace-cross Abstract: Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments when support is relevant, a challenge we term the diary paradox.

By Zhiyuan Wang, Subigya Nepal, Ariful Islam, Indrajeet Ghosh, Xinyu Chen, Katharine E. Daniel, Laura E. Barnes, Philip Chow
arXiv Computation and Language
Sep 7

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.

By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
arXiv AI
Jun 18

Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep

arXiv:2606. 18596v1 Announce Type: cross Abstract: Sleep diaries are central to behavioral sleep medicine and cognitive behavioral therapy for insomnia, yet daily completion is difficult to sustain, and static forms often provide limited context for interpreting night-to-night sleep variation.

By Amama Mahmood, Bokyung Kim, Honghao Zhao, Molly E. Atwood, Luis F. Buenaver, Michael T. Smith, Chien-Ming Huang
arXiv Machine Learning
Sep 14

On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health

The paper investigates on-device language models (ODLMs) for predicting stress in a mobile health context, focusing on privacy-preserving, cloud-independent inference. Using zero‑shot prompting, the authors evaluate ODLMs across multimodal data—objective sensor features and subjective self‑reports—measuring predictive accuracy, latency, and throughput. Results indicate that sensor features slightly outperform self‑reports, and that lightweight sub‑2B models deliver low latency with predictable resource usage, underscoring both the potential and practical limits of ODLMs for mobile mental health.

By Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker
arXiv Machine Learning
Sep 1

Learning Human Health and Diseases from 24-hour Wrist Movement

The paper introduces Sensori, a self‑supervised foundation model that learns health representations from 24‑hour raw tri‑axial wrist movement data. Trained on 122,640 participants across the UK, China, and the US, Sensori captures diverse movement behaviours, demographics, health axes, and physical function. In independent cohorts, the model improved disease classification for 52 of 102 conditions and incident disease risk prediction for 26 of 87 conditions, especially for neurological and psychiatric disorders.

By Yong Wang, Dylan McGagh, Katya Broomberg, Zizheng Zhang, Jonathan Carter, Junayed Naushad, Laura Brocklebank, Yang Sun, George Nicholson, Dianjianyi Sun, Canqing Yu, Jun Lv, Maxim Barnard, Hubert Lam, Andrew Steptoe, David W. Eyre, Liming Li, Zhengming Chen, Naomi Wray, Spiros Denaxas, Gary S. Collins, Huaidong Du, Aiden Doherty, Hang Yuan
arXiv Computation and Language
Aug 28

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.

By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv AI
Jun 15

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

arXiv:2606. 14604v1 Announce Type: cross Abstract: Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking.

By Pavlos Nicolaou, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios