arXiv:2606. 18147v1 Announce Type: new Abstract: Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians.
By Yuwei Zhang, Tong Xia, Bianca Emmerich, Yu Yvonne Wu, Dimitris Spathis, Xin Liu, Daniel McDuff, Cecilia Mascolo
BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.
By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
PhysioBench is a unified benchmark for physiological signal question answering that consolidates annotations from 22 public datasets into 61.4 million question‑answer pairs spanning 30 tasks. Each pair is linked to a specific signal segment and traceable to its source annotation. The benchmark evaluates 21 models—including large language models, vision‑language models, time‑series language models, and physiological signal foundation models—under three settings, revealing that no model consistently excels across all modalities and tasks, and that performance is sensitive to question wording.
By Mengxuan Li, Junfa Chen, Jinze Xia, Yundan Chen, Lixin Fan, Ke Liu, Keyue Shi, Haishuai Wang
arXiv:2601.13880v2 Announce Type: replace
Abstract: Personalized lifestyle health analysis requires long-horizon, multi-dimensional reasoning over heterogeneous lifestyle signals, and recent advances...
By Ye Tian, Zihao Wang, Onat Gungor, Xiaoran Fan, Tajana Rosing
arXiv:2609.06976v1 Announce Type: new
Abstract: As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients...
By Yuchen Niu, Yanan Ma, Srinivasan Nandakumar, Maolin Chen, Viktor Schlegel, Kexin Wei, Ling Cheng, Anna Bird, Anil Anthony Bharath, Siew-Kei Lam
arXiv:2603. 06638v3 Announce Type: replace-cross Abstract: The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning.
By Sirui Li, Shuhan Xiao, Mihir Joshi, Ahmed Metwally, Daniel McDuff, Wei Wang, Yuzhe Yang
The paper introduces Affective Agent, a three‑layer reference architecture designed for on‑device personalized intervention reasoning in wearable systems. It integrates a compact sub‑billion‑parameter language model with physiological data, context, and user history to determine when and how to intervene, all without cloud support or per‑user retraining. The architecture’s perception, personalization, and reasoning layers adapt through host‑managed structured memory evolution, and evaluation on simulated indoor environmental quality scenarios shows that memory‑driven personalization and two‑pass reasoning enhance intervention decisions.
By Reina Mun, Zishen Wan, Vijay Janapa Reddi
arXiv:2606. 02802v1 Announce Type: new Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs).
By Bo-Hong Wang, Baicheng Peng, Ruilin Wang, Jun Bai, Ziyang Song, Yue Li
arXiv:2605. 22759v2 Announce Type: replace Abstract: While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challenging.
By Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari, Samy Abdel-Ghaffar, Marius Guerard, Kara Vaillancourt, Zhihan Zhang, Jake Garrison, Levi Albuquerque, Dimitris Spathis, Hong Yu, Hamid Palangi, Xuhai "Orson" Xu, David G. T. Barrett, Joseph Breda, Jed McGiffin, Yubin Kim, Yuwei Zhang, Naghmeh Rezaei, Samuel Solomon, Karan Ahuja, Tim Althoff, Jake Sunshine, Ming-Zher Poh, Benjamin Yetton, Ari Winbush, Nicholas B. Allen, James M. Rehg, Isaac Galatzer-Levy, Yun Liu, John Hernandez, Anupam Pathak, Conor Heneghan, Yuzhe Yang, Ahmed A. Metwally, Pushmeet Kohli, Mark Malhotra, Shwetak Patel, Xin Liu, Daniel McDuff
arXiv:2607. 13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation.
By Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo
arXiv:2607. 09322v1 Announce Type: new Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making.
By Yanzhen Chen, Zihan Xu, Xiaocheng Zhang, Zhiting Fan, Weiqi Zhai, Hongxia Xu, Zuozhu Liu