BALMS is a benchmark for evaluating large language model (LLM) agents that analyze longitudinal wearable data to predict mental‑health wellbeing scores and generate evidence‑grounded rationales. It covers three real‑world datasets, two task families (score prediction and rationale generation), and tests five LLM backbones across open‑ and closed‑source paradigms. The study finds that zero‑shot agents rarely beat a simple mean baseline, and while chain‑of‑thought prompting helps reasoning, it does not ensure temporal grounding or numerical accuracy.
By Yu Yvonne Wu, Arvind Pillai, Yuliang Chen, Yuwei Zhang, Sudarshan Regmi, Tess Z. Griffin, Michael V. Heinz, Lisa A. Marsch, Nicholas C. Jacobson, Andrew Campbell
arXiv:2606. 02812v1 Announce Type: new Abstract: Modeling patient trajectories from longitudinal electronic health records (EHRs) requires reasoning over sparse, noisy, and long-context multimodal sequences.
By Sihang Zeng, Matthew Thompson, Ruth Etzioni, Meliha Yetisgen
arXiv:2607. 13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation.
By Haoran Li, Jiebi Deng, Tong Jin, Jinghong Han, Yuxin Wang, Zexin Wang, Qingyi Si, Weikang Gong, Xiahai Zhuang, Jia You, Wei Cheng, Jianfeng Feng, Hongcheng Guo
arXiv:2606. 17767v1 Announce Type: cross Abstract: Personal health data from wearables are typically presented through dashboards of charts and summary statistics, requiring users to actively interpret patterns and implications.
By Nikola Kovacevic, Bastien Husler, Di Zhuang, Rafael Wampfler, Barbara Solenthaler
arXiv:2606. 15504v1 Announce Type: new Abstract: In recent years, the advances of large language models and autonomous agents have revolutionized the healthcare field, facilitating diagnosis and improving treatment results.
By Qianxue Zhang, Yiming Ren, Shihuan Qin, Xiao Zhang, Liao Zhang, Jinyang Huang, Zhengliang Liu, Chenbin Liu, Hongying Feng, Jingyuan Chen, Yuzhen Ding, Weihang You, Hanqi Jiang, Yi Pan, Yifan Zhou, Junhao Chen, Lifeng Chen, Wei Liu, Tianming Liu, Zengren Zhao, Lian Zhang
arXiv:2605. 29483v2 Announce Type: replace Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-specific prediction pipelines or reactive question answering over static summaries.
By Di Zhu, Yu Yvonne Wu, Hong Jia, Aaqib Saeed, Vassilis Kostakos, Ting Dang
EmoMed is a multimodal medical consultation agent that tailors its responses to users' emotional states—such as anxiety, confusion, or urgency—while preserving clinical accuracy. It processes text and medical images, detects affect indicators, and adjusts tone, structure, and detail accordingly. The system ensures factual reliability through a dual retrieval mechanism that combines web-based fact‑checking with an API‑connected, continuously updated medical knowledge base, and it has been evaluated across seven state‑of‑the‑art language models using comprehensive metrics, showing that emotionally adaptive responses outperform neutral baselines without sacrificing accuracy.
By Ivan Nasonov, Nikita Glazkov, Ivan Makovetskiy, Mikhail Mozikov, Daniil Sukhorukov, Andrey Savchenko, Ilya Makarov
arXiv:2608. 10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication.
By Qianggang Ding, Xingyao Wang, Rui Feng, Zhibin Wang, Feixiang Wang, Kelong Mao, Hao Sun, Zhiyao Luo, Jiankai Tang, Lei Li, Jiadong Guo, Minheng Ni, Weicong Lin, Chenxi Yang, Hongxiang Gao, Zhenghua Chen, Yang Bai, Min Wu, Jun Cheng, Huazhu Fu, Dacheng Tao, Bang Liu
MERID is a framework that uses recursive self‑improvement agents to autonomously develop multimodal pipelines for detecting major depressive disorder. It aligns multimodal records with depression targets, jointly modifies representations, fusion, and predictors, and guides revisions through evidence‑guided evolution to validate improvements before inheritance. Experiments on depression benchmarks show MERID outperforms existing multimodal and agent‑based baselines, especially highlighting the importance of acoustic and linguistic cues.
By Lei Liu, Zhaokang Liang, Qingcheng Zeng, Chenda Duan, Lu Mi, Zhen Tan, Tianyu Liu
arXiv:2608. 06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.
By Abdulkadir K\"ul\c{c}e, Alihan Esen, Ca\u{g}la Fikir, Berke Kurt, Kuzey Arar, G\"okhan Ercan, Faik Boray Tek
arXiv:2606. 04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agent have become essential.
By Manvendra Modgil
arXiv:2608. 06380v1 Announce Type: cross Abstract: Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-mediated inflammatory diseases (IMID).
By Julian Fierrez, Alejandro Pe\~na, Aythami Morales, Ruben Tolosana, Ruben Vera-Rodriguez, Meenakshi Chatterjee, Ahmaniemi Teemu, Wan-Fai Ng, Walter Maetzler, Nikolay V. Manyakov, Jennifer Kudelka, Ralf Reilmann, C. Janneke van der Woude, Kristen Davies, Victoria Macrae, IDEA-FAST Consortium