WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.
By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
arXiv:2606. 18147v1 Announce Type: new Abstract: Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians.
By Yuwei Zhang, Tong Xia, Bianca Emmerich, Yu Yvonne Wu, Dimitris Spathis, Xin Liu, Daniel McDuff, Cecilia Mascolo
Wearable and mobile sensing technologies have demonstrated strong potential for health inference; however, most sensor models are designed for specific disease types, limiting their transferability across different health risks. Wearable foundation models offer a more generalizable approach in diverse health risk types.
arXiv:2605. 29483v2 Announce Type: replace Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines or reactive question answering over static summaries.
By Di Zhu, Yu Yvonne Wu, Hong Jia, Aaqib Saeed, Vassilis Kostakos, Ting Dang
The paper introduces Affective Agent, a three‑layer reference architecture designed for on‑device personalized intervention reasoning in wearable systems. It integrates a compact sub‑billion‑parameter language model with physiological data, context, and user history to determine when and how to intervene, all without cloud support or per‑user retraining. The architecture’s perception, personalization, and reasoning layers adapt through host‑managed structured memory evolution, and evaluation on simulated indoor environmental quality scenarios shows that memory‑driven personalization and two‑pass reasoning enhance intervention decisions.
By Reina Mun, Zishen Wan, Vijay Janapa Reddi
arXiv:2607. 13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation.
By Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo
arXiv:2607. 16235v1 Announce Type: cross Abstract: Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching.
By Narayan Schuetz, Yuze Bai, Lianggang Pan, Edgar Eggert, Favour Nerrise, Juan Delgado-SanMartin, Max Rosenblattl, Milana Gurbanova, Mohammad Asadi, Anders Johnson, Paul Schmiedmayer, Dennis Wang, Allan Lawrie, Daniel Seung Kim, Xin Liu, Akshay Paruchuri, Ehsan Adeli, Euan Ashley, Kelly W. Zhang
arXiv:2607. 06954v1 Announce Type: new Abstract: Wearable and mobile sensing technologies have demonstrated strong potential for health inference; however, most sensor models are designed for specific disease types, limiting their transferability across different health risks.
By Zhenghuang Wu, Yuyao Zhu, Songlin Xu
arXiv:2605. 22759v2 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging.
By Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
arXiv:2608. 03251v1 Announce Type: cross Abstract: Commercial wearable devices continuously capture rich physiological data (e.
By Esther Brown, Karis Moon, Victoria Dean, Finale Doshi-Velez
arXiv:2512. 08211v2 Announce Type: replace Abstract: Large language models (LLMs) are moving from cloud-centric services toward on-device embedded AI, where models interact with private, longitudinal signals sensed from users and their physical environments.
By Jiaxiang Geng, Lunyu Zhao, Yiyi Lu, Bing Luo
The paper investigates on-device language models (ODLMs) for predicting stress in a mobile health context, focusing on privacy-preserving, cloud-independent inference. Using zero‑shot prompting, the authors evaluate ODLMs across multimodal data—objective sensor features and subjective self‑reports—measuring predictive accuracy, latency, and throughput. Results indicate that sensor features slightly outperform self‑reports, and that lightweight sub‑2B models deliver low latency with predictable resource usage, underscoring both the potential and practical limits of ODLMs for mobile mental health.
By Ibukunoluwa Soyebo, Alyssa Donawa, Rodrigo Aguilar Barrios, Brice Patchou, Corey E. Baker