arXiv Machine Learning

Explainable Prediction from Mobile Sensing Data through LLM-guided Concept Integration

The paper introduces a Concept-Integrated Transformer (CIT) that uses a pretrained large language model to generate concept abnormality targets with confidence weights, eliminating the need for manual concept annotation. CIT is applied to mobile sensing data from two longitudinal datasets, achieving the highest F1 score on the AFFECT dataset (0.756) and tying for the highest on a PHQ-9 dataset (0.765). The model’s learned concept scores reveal interpretable behavioral and physiological patterns, such as differences in sleep quantity and quality between high and low negative affect groups.

arXiv AI
Jun 2

Towards a General Intelligence and Interface for Wearable Health Data

arXiv:2605. 22759v2 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging.

By Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
arXiv Computation and Language
Aug 28

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.

By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv Machine Learning
Aug 27

Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions

The study introduces a clinician‑in‑the‑loop benchmark to assess whether large language models can generate evidence‑grounded Brief Hierarchical Taxonomy of Psychopathology (B‑HiTOP) item profiles from multimodal data, including passive sensing, ecological momentary assessment, and questionnaires. Using the GLOBEM dataset, the authors create 14,592 participant‑day instances aligned to 29 B‑HiTOP items across five spectra, and evaluate evidence compatibility rather than diagnostic accuracy. Two‑stage prediction improves compatibility for EMA and questionnaire evidence but reduces it for passive sensing and combined evidence, yielding more conservative score distributions across models, spectra, and evidence settings.

By Xiyun Hu, Xiangyuan Xue, Yuting Lyu, Hanya Shao, Jingping Nie
arXiv Machine Learning
Sep 14

On-Device Language Models for Privacy-Preserving Stress Prediction: A Multimodal Evaluation on Mobile Health

The paper investigates on-device language models (ODLMs) for predicting stress in a mobile health context, focusing on privacy-preserving, cloud-independent inference. Using zero‑shot prompting, the authors evaluate ODLMs across multimodal data—objective sensor features and subjective self‑reports—measuring predictive accuracy, latency, and throughput. Results indicate that sensor features slightly outperform self‑reports, and that lightweight sub‑2B models deliver low latency with predictable resource usage, underscoring both the potential and practical limits of ODLMs for mobile mental health.

By Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker
arXiv Machine Learning
Jul 9

Counterfactual Modeling with Fine-Tuned LLMs for Health Intervention Design and Sensor Data Augmentation

arXiv:2601. 14590v3 Announce Type: replace Abstract: Counterfactual explanations (CFEs) provide human-centric interpretability by identifying the minimal, actionable changes required to alter a machine learning model's prediction.

By Shovito Barua Soumma, Asiful Arefeen, Stephanie M. Carpenter, Melanie Hingle, Hassan Ghasemzadeh
arXiv AI
Aug 18

Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection

arXiv:2608. 14675v1 Announce Type: cross Abstract: While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as photoplethysmography (PPG), its suitability for highly subjective tasks remains unproven.

By Dominika Kunc, Przemys{\l}aw Kazienko, Stanis{\l}aw Saganowski
arXiv AI
Jun 15

A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health

arXiv:2606. 14604v1 Announce Type: cross Abstract: Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comparisons of modern forecasting architectures for these data are lacking.

By Pavlos Nicolaou, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios
arXiv Machine Learning
Jul 30

HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring

arXiv:2509. 07260v5 Announce Type: replace-cross Abstract: Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and ultimately improving individuals' quality of life.

By Xin Wang, Ting Dang, Xinyu Zhang, Vassilis Kostakos, Michael J. Witbrock, Hong Jia