arXiv AI By Di Zhu, Yu Yvonne Wu, Hong Jia, Aaqib Saeed, Vassilis Kostakos, Ting Dang

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Read the original on arXiv AI →

arXiv:2605. 29483v2 Announce Type: replace Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines or reactive question answering over static summaries.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing

BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.

By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv Computation and Language
Sep 7

WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data

WearableQA is a new benchmark that tests AI systems on health reasoning using real-world wearable data from 200 users, each with up to 500 days of daily measurements. It contains 4,084 ten‑option multiple‑choice questions derived from wearable time series, blood biomarkers, and demographics, and is organized into 16 question types that distinguish data‑driven computation from physiological interpretation and single‑signal from cross‑signal reasoning. Evaluation of 14 large language models shows wide performance gaps, indicating that the benchmark remains challenging and useful for diagnosing model capabilities.

By Ji Soo Lee, Xilun Chen, Pierce Chuang, Ashish Shenoy, Jason Wei, Dohwan Ko, Hyunwoo J. Kim, Benoit Corda
arXiv Machine Learning
Sep 21

PhysioBench: A Unified Benchmark for Physiological Signal Question Answering

PhysioBench is a unified benchmark for physiological signal question answering that consolidates annotations from 22 public datasets into 61.4 million question‑answer pairs spanning 30 tasks. Each pair is linked to a specific signal segment and traceable to its source annotation. The benchmark evaluates 21 models—including large language models, vision‑language models, time‑series language models, and physiological signal foundation models—under three settings, revealing that no model consistently excels across all modalities and tasks, and that performance is sensitive to question wording.

By Mengxuan Li, Junfa Chen, Jinze Xia, Yundan Chen, Lixin Fan, Ke Liu, Keyue Shi, Haishuai Wang
arXiv Machine Learning
Sep 10

HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care

arXiv:2609.06976v1 Announce Type: new Abstract: As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients...

By Yuchen Niu, Yanan Ma, Srinivasan Nandakumar, Maolin Chen, Viktor Schlegel, Kexin Wei, Ling Cheng, Anna Bird, Anil Anthony Bharath, Siew-Kei Lam